The AI Cost Index estimates the levelized cost of delivering the same frontier inference service in 45 countries. Model quality, request profiles, latency requirements and production throughput are held constant. Hardware may differ where sovereignty or export controls require it, but each configuration must provide the same service.
The headline unit is US dollars per million output tokens. Two deployments and three workloads are published separately. They are scenarios, not components of a composite score.
The current results are screening estimates. They are designed to make assumptions comparable and debatable. They are not supplier bids or investment recommendations.
What the index measures
The index asks a narrow question:
What would it cost to produce an equivalent frontier inference service inside each country?
It is not an AI readiness index. It does not award points for universities, patents, regulation or startup activity. It estimates the annualized production cost of compute, electricity, facilities, grid connections and operations.
The denominator is production output measured under the same workload, model-quality and latency requirements in every country. Countries do not receive different acceptance rates. Failed requests and retries consume capacity without producing output, so their effect belongs in measured production throughput rather than a separate country score.
Reference workloads
| Workload | Request profile | Service requirement | Throughput | Average IT power |
|---|---|---|---|---|
| General inference | 8k input, 2k visible output, medium reasoning | TTFT ≤5 seconds; output ≥20 token/s | 3,683 token/s | 98 kW |
| Long-context analysis | 128k input, 4k visible output | TTFT ≤20 seconds; no quality regression | 1,719 token/s | 124 kW |
| Agentic engineering | Multi-turn tools, code and structured output | Task deadline and correctness suite | 777 token/s | 132 kW |
General inference is the default because it is the broadest baseline. It is not assumed to be more economically important than long-context or agentic work.
General-inference throughput is calibrated to Kimi K3 API economics. Kimi publishes prices of $15 per million output tokens, $3 per million cache-miss input tokens and $0.30 per million cache-hit input tokens. It reports a cache-hit rate above 90% for coding workloads. The index uses 90% exactly. With four input tokens per output token, API revenue is therefore:
$15 + 4 × (90% × $0.30 + 10% × $3.00) = $17.28
The calibration assumes Kimi subsidizes no more than 25% of production cost. The maximum implied production cost is therefore:
$17.28 / (1 - 25%) = $23.04 per million output tokens
Solving the China 100 MW reference case for that cost raises general-inference throughput from the original 900 token/s engineering prior to 3,683 token/s per 64-to-72-accelerator service unit. This is aggregate production output across concurrent requests, not single-user generation speed.
The resulting 4.092× efficiency multiplier is also applied to the original long-context and agentic priors. Those two workloads are extrapolations rather than direct API-price calibrations. Because the Kimi calibration uses the maximum permitted subsidy, it is a conservative cost ceiling. A positive provider margin, a cache-hit rate above 90% or a subsidy below 25% would imply higher throughput and lower production cost.
For agentic use, a future release should also publish dollars per 1,000 successfully completed reference tasks. Token cost can otherwise reward shorter output even when fewer tasks are solved.
Deployment models
The two deployments are calculated independently.
| Deployment | Central engineering treatment |
|---|---|
| 100 MW data center | 100 MW at the utility meter; 82 to 93.5 MW commissioned IT capacity after PUE; approximately 577 to 658 service units; 80% utilization; centralized operations |
| On-prem solution | One rack-scale service unit; 250 kW connection envelope; 70% utilization; industrial tariff and on-premises facility allowance |
The 100 MW result is not the on-prem result multiplied by 400. It includes a bulk hardware discount, large-load power contracting, more efficient cooling, centralized staffing and a separate greenfield construction model.
Mathematics
Annual levelized cost is the sum of hardware capital recovery, site and grid capital recovery, electricity and operating expense.
Annual cost = CRF(WACC, 3 years) × landed compute hardware
+ CRF(WACC, asset life) × facility and grid connection
+ delivered electricity
+ support, software, labour and workload overhead
CRF(r,n) = r(1+r)^n / ((1+r)^n - 1)
Annual output is:
service units × output tokens per second × 31,536,000
× utilization × technical availability
The headline cost is:
AI cost = $1,000,000 × annual levelized cost / annual output tokens
The central reference service capacity uses a $4.59 million GB300-class calibration before country multipliers, a three-year hardware life and zero residual value. The model converts euro-denominated inputs to US dollars at 1 euro = $1.1467, the ECB reference rate of 16 July 2026.
Historical backcast
The historical view is a fixed-service counterfactual. It asks what the 2026 frontier inference service would have cost under the economic conditions of each year from 2010 to 2026.
The following quantities remain fixed:
- Model quality, workloads and latency requirements
- Hardware capacity and real purchase price
- Output throughput and utilization
- Hardware life and resilience design
- Country hardware-landing multiplier
- Country PUE and service availability
The following inputs change annually:
- Industrial electricity price
- Cost of capital
- Facility and grid construction cost
- Exchange rates and general inflation
All historical values are reported in constant 2026 US dollars. Holding hardware capability and real purchase price constant prevents semiconductor progress from dominating the economic question. Financing still changes the annualized cost of owning that hardware.
European power prices use the Eurostat nrg_pc_205 500 to 1,999 MWh band, excluding recoverable VAT, averaged across each year’s two semesters. US power prices use the EIA annual industrial average. Both are converted into constant 2026 US dollars. Canada and Mexico use an EIA-anchored North American path. Other non-European countries use explicit regional real-price paths anchored to their 2026 country input and should be treated as modeled estimates.
Annual financing changes use the FRED DGS10 US 10-year Treasury yield. Each country’s 2026 spread is held constant, with additional euro-area sovereign-crisis adjustments for Greece, Portugal, Ireland, Spain, Italy and Cyprus. This is a transparent financing proxy rather than a reconstructed project-finance quote.
Facility and grid capex use the FRED new-industrial-building construction producer price index, deflated by the US consumer price index and normalized to 2026. Operations are held constant in real terms.
The annual conditions index is:
Annual conditions index(c,y)
= fixed hardware annualized at WACC(c,y)
+ electricity(c,y)
+ facility and grid capex(c,y)
+ fixed real operations
------------------------------------------------
fixed annual output
The global historical line is the unweighted median of the 45 country estimates, normalized so that 2026 equals 100. A value of 93 means that the median modeled owner cost was 7% below the 2026 level. A peak identifies unusually expensive ownership conditions, but does not by itself prove an asset-price bubble.
The backcast does not assert that GB300 hardware existed in 2010. It reprices a fixed 2026 basket under historical economic conditions, similar to holding the contents of a price index constant through time.
Country inputs
The ranking incorporates more than electricity price:
- Delivered industrial electricity cost
- Power usage effectiveness
- Legally available hardware and landed hardware premium
- Cost of capital
- Facility construction cost
- Grid connection and reinforcement cost
- Operations and technical labour
- Utilization and residual service availability
EU electricity values use Eurostat non-household observations. The United States uses the EIA industrial average. Other markets use regulator, utility and statistical sources where available, supplemented by explicit country proxies.
Hardware export controls are treated as a cost input, not a reason to create a separate ranking. Each country uses the best legally deployable configuration capable of passing the common workload. Additional accelerators, power, porting work and support are charged to that country.
China uses a Huawei sovereign-system equivalent normalized to the same service target and approximately the same $4.59 million capital calibration. Public price and full-workload throughput evidence is not yet sufficient to verify that assumption, so China receives a Grade D evidence rating.
Reliability and the grid
The facility includes standard UPS and standby generation. National outage minutes are therefore not deducted directly from compute output.
SAIDI measures interruption duration. SAIFI measures interruption frequency. SAIFI influences UPS cycles and generator starts, while SAIDI influences backup-fuel runtime. An investment-grade version should use feeder-level planned and unplanned interruption data, momentary interruptions, voltage quality and a common exceptional-event convention.
The final site model should also replace national electricity proxies with an hourly PPA or utility tariff that includes demand, capacity, balancing and non-recoverable taxes.
Why AI access could matter for GDP
The economic channel is task-level:
Potential productivity effect ≈ task exposure
× adoption and usage intensity
× value created or cost saved per task
Lower inference cost makes more marginal tasks economical and permits greater use. It does not guarantee adoption or productivity. Organizational change, capital access, skills and distribution still determine whether capacity is used effectively.
If mature models converge in quality, inference may begin to resemble electricity, cloud infrastructure or telecommunications. Price, abundance, reliability and distribution would then become national competitiveness variables. If a sharp capability discontinuity appears, control of the leading model could matter more than token economics.
Cost is only one part of national AI access. A future companion index should measure annual inference capacity per worker or per unit of GDP at a common budget.
Evidence and uncertainty
Country inputs receive one of three grades:
| Grade | Interpretation | Country-specific range |
|---|---|---|
| B | Official statistical price or credible regulator or utility tariff | ±12% |
| C | Regional schedule, older observation or modeled national proxy | ±18% |
| D | Material hardware price or performance evidence gap | ±25% |
The general-inference absolute level is price-implied rather than directly benchmarked. Its uncertainty is asymmetric: the 25% maximum-subsidy assumption defines a conservative cost ceiling, while positive provider margin or greater serving efficiency would lower cost. Long-context and agentic throughput retain a common ±35% uncertainty because they inherit the K3 efficiency multiplier without direct workload-specific price calibration. These common uncertainties have less effect on pairwise country comparisons than on absolute dollar values.
Limitations
The release does not yet include binding 100 MW utility offers, individual demand tariffs, hourly hedging, tax structures, grants, carbon pricing, disaster-recovery duplication, heat revenue or workload-specific task success. Country financing, hardware landing and site-capex multipliers remain modeled proxies.
An investment-grade release should obtain hardware quotations, grid-feasibility studies, large-load tariffs or PPA proposals and data-center designs for at least three candidate sites in each country. It should then benchmark the exact serving stack under the published workloads.
Downloads
- Country estimates in CSV
- Complete structured dataset
- Historical estimates in CSV
- Historical estimates in JSON
- Historical annual input table
- Python calculation model
- Python backcast model
Principal sources
- NVIDIA GB300 NVL72 system reference
- Kimi K3 architecture, serving design and API pricing
- European Central Bank exchange rates
- Eurostat non-household electricity prices
- US Energy Information Administration electricity prices
- FRED 10-year Treasury rate
- FRED US consumer price index
- FRED new industrial building construction price index
- International Energy Agency: Energy and AI
- ISO/IEC 30134-2:2026 Power Usage Effectiveness