corpusAI

Index methodology

How every corpusAI index fixing is computed, versioned so a number can be cited.

Also as methodology.md

Version 2026.09.14-2 (adds the EXE executed-price segment; 2026.09.14-1 added the LLM token price index; 2026.09.13-3 added the AGG segment and TensorDock / io.net to MKT; 2026.09.13-2 added Azure hyperscaler legs and the H100NVL code; earlier fixings carry 2026.09.13-1). Every fixing carries this version. A definition below never changes under the same version; a change ships as a new version and, for a changed series, a new ticker.

What the indices measure

The corpusAI compute price indices are public-quote indices: each fixing is computed from prices that anyone could have paid on that day, observed on the provider's own public price feed or, for hyperscaler spot, from the provider's own price-change API. They do not include brokered, private or negotiated transactions. That is a deliberate scope: an agent can act on a public quote.

Unit

USD per GPU-hour. A quote for a unit holding n GPUs is divided by n; a quote already expressed per GPU is used as is. Only USD quotes are used.

GPU models

Quotes are mapped to a normalised model name (/gpu/models). Form factor and memory are distinguished only where they change price: H100 SXM, H100 PCIe, H100 NVL, A100 SXM 80GB, A100 SXM 40GB, A100 PCIe 80GB. Tickers use the model codes H100SXM, H100PCIE, H100NVL, H200, B200, GB200, A100SXM80, A100SXM40, A100PCIE80, L40S, L4, A10, T4, V100, RTX4090, RTX5090, RTX6000ADA, RTXPRO6000, MI300X.

Segments

SegmentTickerMembers (provider/market)
MarketplaceMKTVast.ai on-demand, RunPod community, TensorDock, io.net (regional and network)
Aggregator listingsAGGShadeform, Prime Intellect, Spheron on-demand re-listings, kept apart from NEO so a neo-cloud price is not counted twice
Neo-cloud on-demandNEORunPod secure, Lambda, DataCrunch, Nebius, Crusoe, CoreWeave, Hyperbolic on-demand
InterruptibleINTVast.ai minimum bid, RunPod spot, DataCrunch spot, Nebius preemptible, CoreWeave spot, Crusoe spot
ExecutedEXENosana jobs that actually ran, and Vast.ai offers taken between snapshots (inferred at their last ask)
Hyperscaler on-demandHYPOracle, Vultr, Linode list prices, plus AWS, GCP and Azure on-demand for the mapped instance family ÷ GPUs per node
Hyperscaler spotSPOTAWS, GCP and Azure spot for the mapped instance family ÷ GPUs per node; SPOT.<REGION> for one region

A ticker is CX.<MODEL>.<SEGMENT>, for example CX.H100SXM.NEO, or CX.H100SXM.SPOT.US-EAST-1 for a hyperscaler region. /index/tickers lists every ticker that can currently fix.

Instance mapping for hyperscaler legs

ModelAWSGCPAzureGPUs per node
H100 SXMp5.48xlargea3-highgpu-8gND96isr_H100_v58
H100 NVLNC40ads_H100_v51
H200p5en.48xlarge8
B200p6-b200.48xlarge8
A100 SXM 80GBp4de.24xlargea2-ultragpu-8gND96amsr_A100_v48
A100 SXM 40GBp4d.24xlargea2-highgpu-8gND96asr_A100_v48
A100 PCIe 80GBNC24ads_A100_v41
L40Sg6e.xlarge1
L4g6.xlargeg2-standard-41
A10g5.xlargeNV36ads_A10_v51
T4g4dn.xlargeNC4as_T4_v31
V100p3.2xlarge1
MI300XND96isr_MI300X_v58

The mapped instance is the smallest unit whose price is dominated by the GPU. For 8-GPU nodes the per-GPU price includes the node's CPU, memory and network; that is disclosed, not corrected.

Eligibility

A quote enters a fixing when all of the following hold:

Hyperscaler spot legs have one extra rule: Azure publishes token spot prices (fractions of a cent per node-hour) for GPU SKUs it has no spot capacity for. A leg below 3% of the instance's on-demand list price, or below $0.02 per GPU-hour where no list price is known, is a placeholder, not a price anyone can pay, and is excluded from fixings and histories.

Vast.ai queries return the cheapest 64 offers per model, so the marketplace segment is a cheapest-offers index, not a whole-book index; the offer count is reported as n on every fixing.

Aggregation

The fixing is the median of eligible quotes in the segment. The lowest and highest eligible quote and the count are reported beside it. There is no capacity weighting, because capacity is not observable on public feeds.

Hyperscaler spot legs are the time-weighted daily average price of the mapped instance (a price counts for as long as it was in force, per availability zone, zones collapsed by median), divided by GPUs per node. The SPOT composite is the median across regions; SPOT.<REGION> is one region. spread_vs_neo on a SPOT fixing is the composite minus the NEO fixing of the same day.

Executed prices

EXE is the only segment whose members are prices somebody paid rather than asked. Two sources:

The fixing is the median across both, with n and members showing how many rows each contributed, so a day with two fills is visibly a day with two fills.

Capacity stress

/capacity/* is a derived signal, not a price. For a GPU model and day it combines, as a plain mean of whichever components exist:

ComponentInputStress
availabilityShadeform and Lambda on-demand listings with an availability flag100 × (1 − available ÷ listed)
ionetio.net regional SKUs: deployable units ÷ total units100 × (1 − share)
supplyVast.ai on-demand offers seen today ÷ trailing 30-day median100 × max(0, 1 − ratio)
spot_ratiohyperscaler SPOT fixing (previous finished day) ÷ HYP fixing100 × clamp((ratio − 0.3) ÷ 0.6)

0 means slack, 100 means tight. Components and their inputs are returned with every score; the board lists only models with at least two components.

LLM token price index

Rows are list prices in USD per million tokens from four public sources collected once a day: the OpenRouter model catalog (routed price per model), OpenRouter's per-host endpoints for the index models (the same model as priced by each hosting provider), the LiteLLM price table (MIT, the SDK's view of every provider's list price, chat models only), and DeepInfra's model list. Model names are reduced to a canonical slug (vendor and host prefixes, OpenRouter variant suffixes and Bedrock version tails removed) so the same model lines up across sources.

A fixing is the median across every source and host row for the model that day at the standard tier: rows tagged as a variant (batch, free, extended, thinking and similar) and rows with a zero price are excluded. .IN and .OUT are the input and output medians; .BLEND is (3 × input + output) / 4, a 3:1 input:output mix. CX.TOK.FRONTIER.<side> is the median across a fixed basket of frontier models (listed by /tokens/models); the basket changes only with a methodology version. n counts the rows behind a fixing and members names them, so a single-host proprietary model is visibly a single-source number.

Term structure

/term and /term/instance are not index fixings; they are ladders of current list prices. On-demand and committed terms come from the provider price lists we collect (AWS Price List API weekly, Azure Retail Prices API daily, GCP Cloud Billing Catalog daily as vCPU × core + GiB × RAM + GPU SKUs); a committed term is expressed as effective hourly USD with any upfront amortised evenly over the term. The spot point is the current median across zones from our own events. Per-GPU curves apply the instance mapping above and take the median across regions. Neo-cloud commitment prices are not published as numbers and are excluded.

Timing

Quote snapshots are taken once per UTC day from each provider's public feed; the fixing day is the snapshot day. Hyperscaler spot fixings cover the full UTC day and are final once the day has ended, so the latest SPOT fixing is yesterday's.

History and sources

Quote-based segments start 2026-09-13. Hyperscaler spot legs reach back to the start of the underlying spot series: measured daily data from 2022-05-31 for AWS (Pauley dataset, then SpotLake TITANS daily averages, then our own price-change events from 2026-06-06), 2024 for GCP, and from the first Azure collection of each GPU SKU (2026-09-13 for most GPU SKUs). Each daily value carries its source (pauley, titans, direct); only direct rows are computed from full price-change events.

Revisions

A published fixing is final. Late or corrected data is reflected in the next fixing, never by restating an earlier one. Parser changes on a provider feed are noted in the changelog; a change that alters what a segment measures is a new methodology version.

Known limitations