InferenceIndexerInferenceIndexer.ai
/
ModelsProvidersAPIEmbedFor agentsHarnessesMethodologyAboutLoginSign Up

SIT Methodology

How the recommendation engine ranks models, how the Standard Inference Token price is defined, calculated, and governed.

Version 0.4, last updated August 6, 2026

1.Overview

The Standard Inference Token (SIT) tracks the marginal cost of producing AI inference tokens at a defined quality standard. The SIT Token Price Index (TPI, formerly SIT-Composite) tracks the market price of one million GPT-4-Turbo-equivalent inference tokens (1 SIT), the commodity unit for AI compute.

The SIT serves four purposes:

  • Price comparison across models on a like-for-like basis
  • Composite indices that track inference price movements over time
  • A reference price that futures contracts can settle against
  • Benchmarking: "am I paying above or below market rate?"

InferenceIndexer is an independent verification and recommendation service for AI inference. We do not provide inference services, do not route API calls, and do not take positions in any inference derivatives market. All data sources are public and verifiable.

2.Definition

2.1Unit

1 SIT = 1 million tokens of inference at a defined quality standard.

"Tokens" refers to the standard industry unit of LLM inference, as measured by each provider's tokenizer. While tokenizers differ between models, the "million tokens" convention is universally adopted and provides sufficient standardization for pricing purposes.

2.2Pricing Components

Every SIT-eligible model has three published prices:

ComponentDefinition
Input priceCost per million input (prompt) tokens
Output priceCost per million output (completion) tokens
Blended priceWeighted average: 40% input + 60% output

Blended price formula:

blended_price = (0.4 × input_price) + (0.6 × output_price)

The 60% output weighting reflects production workloads where output tokens exceed input tokens (generation, coding, summarization). This ratio will be refined as real-world usage data becomes available.

Note: Artificial Analysis uses a different blend (70% cached input, 20% uncached input, 10% output) which assumes heavy prompt caching. Our 40/60 blend assumes no caching, reflecting most production workloads. We also publish a SIT-Cached variant using the 7:2:1 ratio for cached workloads.

2.3Median Pricing (Multi-Provider Models)

Many open-source models (e.g. GLM-5.2, Llama 4, DeepSeek V4) are hosted by multiple inference providers at different prices. For these models, InferenceIndexer computes the median blended price across all available providers, rather than using a single source.

Median price formula:

model_price = median(blended_price_1, blended_price_2, ..., blended_price_n)

The median is used instead of the mean or cheapest price because it is:

  • Robust to outliers: A single provider charging 10x the market rate does not skew the index.
  • Not gameable: A provider cannot manipulate the headline price by temporarily dropping to $0.01.
  • Realistic: The median represents what a user would realistically pay, with half the providers charging more and half charging less.

For models with only one provider, the single available price is used. The number of providers (Sources) is displayed for each model in the table and on the model detail page, where a full provider comparison table shows all available prices.

Provider endpoints are fetched daily from OpenRouter's /api/v1/models/{id}/endpoints endpoint, with direct provider APIs polled hourly. Hourly pipeline runs blend both sources and use the cached median from the most recent fetch.

3.Quality Tiers

Models are grouped into quality tiers based on demonstrated capability, using the Artificial Analysis Intelligence Index as an independent third-party benchmark. Tier boundaries are set as percentiles of the scored model population (top 10% = Frontier, next 20% = Standard, next 30% = Budget, remainder = Micro). Because the cutoffs are relative rather than fixed scores, a benchmark rebase by Artificial Analysis (e.g. Index v4.2 in September 2026) relabels no one: every model's percentile rank is unchanged by a rescale.

TierPopulation shareDescription
SIT-FrontierTop 10% of scored modelsTop-tier models from frontier labs
SIT-StandardP70-P90Mid-tier production models
SIT-BudgetP40-P70Low-cost models for high-volume tasks
SIT-MicroBelow P40 / unscoredUltra-cheap models for simple tasks

Historical note (September 2026): before this change, tiers used fixed AA score thresholds (Frontier >= 50, Standard >= 30, Budget >= 15). When Artificial Analysis shipped Index v4.2 (Sep 4) and v4.3 (Sep 7) within one week, those fixed cutoffs reclassified roughly half the leaderboard twice. Percentile tiers replaced them on Sep 11, 2026.

4.Recommendation Ranking

The recommendation engine turns verified prices and quality scores into ranked, constraint-aware recommendations. It is fully deterministic: the same constraints always produce the same ranking, and no language model is involved anywhere in the ranking path. This section documents exactly how a recommendation is produced; every claim on the homepage reduces to the rules below.

4.1Eligibility

A model can appear in a recommendation only if all of the following hold. Models that fail any test are excluded before ranking, never shown with an "N/A" placeholder:

  • The model is active and has a verified blended price above zero
  • The model has a Cost/IQ value, which requires an Artificial Analysis Intelligence Index score (the AA ≥ 35 pipeline gate applies to having a Cost/IQ at all)
  • The model is not a batch variant
  • The model matches the requested modality. "text" includes vision-capable models (text+image→text), since a vision-capable model can perform text-only work

4.2Hard constraints

Callers filter the eligible pool with hard constraints. A model that fails a hard constraint is excluded, regardless of how well it would rank:

ConstraintMeaning
budget_max_usd_per_mBlended price per million tokens must not exceed this value
context_minContext window must be at least this many tokens
zdrProvider must claim zero data retention (provider-stated, not II-verified)
eu_sovereignProvider must claim EU sovereignty (provider-stated, not II-verified)
reasoningInclude or exclude reasoning models
aa_minArtificial Analysis Intelligence Index must be at least this value
providersRestrict to named endpoint providers

At least one constraint is required; unconstrained browsing is served by the model table instead. Privacy and security constraints are matched on provider statements today, not on II verification. Results say so; see the attestation table on the homepage.

4.3Ranking metric: Cost/IQ

Eligible models are ranked by Cost/IQ ascending. The formula (unchanged from v0.2, August 2026):

Cost/IQ = blended price × (40 / AA Intelligence Index score)

where the blended price is 0.4 × input price + 0.6 × output price. Lower Cost/IQ is better value: it is the verified price of a million tokens normalised per unit of demonstrated intelligence. Cost/IQ is a price-efficiency sort, not a task-fitness score. Because the formula divides by the AA score, a model with a low score and a very low price can rank first; the response carries each model's AA score so this is always visible.

4.4Task adjustments (use_case)

An optional use_case hint applies a deterministic multiplicative adjustment to Cost/IQ after the base ranking. Each profile encodes an AA floor and a reasoning preference:

ProfileAA floorReasoning preferenceRationale
support15AvoidHigh-volume drafting: precision and cost matter more than peak intelligence
volume10AvoidHigh-volume generation: cost dominates
extraction20AvoidStructured output: reasoning adds latency without schema accuracy gains
summarization20AvoidLong context and cost efficiency matter most
coding30PreferredReasoning and higher intelligence correlate with patch quality
research35PreferredReasoning depth and large context matter most

Penalties scale with the shortfall below the profile's AA floor; a reasoning-model penalty of 1.25 reflects that thinking tokens are not reflected in listed prices. Adjustments reorder near-ties; they never override genuine price/quality dominance. Every result shows both the raw and the adjusted basis in its "why" string.

4.5Callable-first demotion

Models with a hand-verified endpoint recipe (a tested provider base_url plus the model's native ID on that provider) rank above models without one. Demoted results carry an explicit caveat and are never silently hidden. The prefer_callable flag (default on) controls this.

4.6Freshness and honesty surface

Every recommendation carries: per-field as-of timestamps, a freshness block with the published SLA (prices 6 hours, AA scores 7 days), the runner-up models with the reason each lost, and the ranking basis in machine-readable form. Queries are never stored: the anonymous demand counter records constraint tuples only (budget, context, flags), never query content.

4.7What is NOT ranked

Latency and uptime are not scored: live probe coverage is too thin to rank on honestly (see the quality-monitoring roadmap). Security posture is not scored: no verification capability exists yet. Privacy is matched on provider statements only. When any of these changes, this section changes with them, before the homepage claims them.

5.Index Calculation

5.1Tier Indices

Each quality tier has its own index, tracking the median blended price per million tokens across all models in that tier. The SIT TPI uses equal weight per provider, capped at 30% (see Section 4.3):

IndexWhat it tracks
SIT-FrontierMedian blended price of all Frontier-tier models
SIT-StandardMedian blended price of all Standard-tier models
SIT-BudgetMedian blended price of all Budget-tier models
SIT TPI (Token Price Index)Equal-weighted mean of cheapest SIT-qualified model per provider, 30% cap (headline price for 1 SIT)
SIT-SpreadFrontier price minus Budget price

5.2Quality-Adjusted Price (Cost / IQ)

The Quality-Adjusted Price is not a transactional price. It is a normalized index value for cross-model comparison. The actual price you pay a provider is the Blended Price. The Quality-Adjusted Price normalizes that price for intelligence so models of different capability levels can be compared on a like-for-like basis.

The quality adjustment uses a transparent benchmark ratio:

  • Quality gate. Only models scoring at or above the GPT-4-Turbo baseline (AA Intelligence Index >= 35) are included in the SIT TPI basket. Models below this threshold are tracked but excluded from the headline number.
  • Intelligence adjustment. Prices are adjusted by the ratio of the GPT-4-Turbo reference score (40) to the model's own Artificial Analysis Intelligence Index score. A model scoring higher than GPT-4-Turbo will have a lower adjusted price (cheaper per unit of intelligence). Lower is better.

The formula:

Cost / IQ = Blended Price × (40 / AA Intelligence Score)
  Lower = cheaper per unit of intelligence
  Comparable across ALL models

SIT Token Price Index (TPI) = Σ(w_p × min_adjusted_price_p)
  for each provider, cheapest SIT-qualified model
  w_p = equal weight per provider, capped at 30%
  eligibility: AA score in top 40% of scored models (P60+)

Where:

  • Blended Price = 0.4 × input + 0.6 × output (per million tokens)
  • 40 = GPT-4-Turbo (Jan 2024) reference score on AA Intelligence Index v4.1. Models scoring 40 are at GPT-4-Turbo parity. The reference is based on the benchmark thresholds defined in the SIT standard (MMLU >= 86%, HumanEval >= 67%, GSM8K >= 92%).
  • AA Intelligence Index = Artificial Analysis Intelligence Index v4.1, an independent third-party benchmark

Lower Cost / IQ = cheaper per unit of intelligence. Cost / IQ is an absolute measure, not relative to any tier median. A model at $0.25/M Cost / IQ is cheaper per unit of intelligence than a model at $1.50/M Cost / IQ, regardless of tier. Models without an AA Intelligence Index score do not receive a Cost / IQ and are excluded from the composite basket.

5.3Provider Equal Weighting (TPI)

The SIT TPI uses a provider equal weighting: each provider contributes their cheapest SIT-qualified model to the basket, and every provider carries equal weight (capped at 30% of total weight). This is the Token Price Index (TPI) methodology from the SIT paper (arXiv:2603.21690).

Equal weight per provider prevents any single provider from dominating the index regardless of how many models they host. The 30% cap prevents large providers from overwhelming the index as more providers are added.

For each provider: cheapest SIT-qualified model's Cost / IQ

TPI = Σ(w_p × min_adjusted_price_p)
  w_p = equal weight per provider, capped at 30%
  eligibility: AA score in top 40% of scored models (P60+)

Per-tier indices (Frontier, Standard, Budget, Micro) use a simple median across all models in that tier. This answers: "What does a typical model in this tier cost?" The SIT TPI uses provider equal weighting to answer: "What does the market charge for GPT-4-equivalent inference?"

5.4Calculation Frequency

FrequencyWhat happens
HourlyPull pricing from all sources, update database
DailyCalculate SIT indices, publish at 00:00 UTC
WeeklyRefresh usage weights from OpenRouter Rankings (Mondays 06:00 UTC)
MonthlyReview tier composition, add/remove models

5.5Base Date and Rebaselining

  • Current base date: September 4, 2026 (SIT TPI = 1000 at this date). Charts mark era breaks where the underlying Artificial Analysis Index version changed.
  • Prior base: August 4, 2026 (superseded Sep 11, 2026).
  • Rebaselining on methodology changes. All rebaselining events are published with full explanation. The September 2026 rebase was forced by two Artificial Analysis Index updates in one week (v4.2 on Sep 4, v4.3 on Sep 7), which changed the composition of the index basket without any real prices moving.

6.SIT Variants

Inference is not a single homogeneous commodity. The SIT supports attribute-based filtering, similar to how CoinMarketCap filters by category (DeFi, Layer 1, etc.).

SIT VariantFilter
SIT-CompositeAll models (headline number)
SIT-FrontierAA Index >= 50
SIT-StandardAA Index 30–49
SIT-BudgetAA Index 15–29
SIT-EU-SovereignEU-hosted only (jurisdiction)
SIT-ZDRZero data retention guaranteed (provider does not store or train on inputs)
SIT-OpenOpen weights models only
SIT-ProprietaryProprietary models only
SIT-CachedWith prompt caching applied (7:2:1 blend)

The SIT TPI is always the headline number. Variant indices allow users to track specific segments of the inference market.

ZDR and EU Infra status: Provider classifications for SIT-ZDR and SIT-EU-Sovereign are based on publicly available provider documentation as of August 2026. ZDR (Zero Data Retention) includes providers that do not store or train on user inputs by default on their standard API. EU-Sovereign includes providers domiciled in the EU/EEA that are not subject to the US CLOUD Act. Many providers offer ZDR or EU hosting only on enterprise plans; these are not classified as ZDR or EU-Sovereign here. Provider policies change; classifications are reviewed quarterly.

7.Data Sources

7.1Primary Sources

All providers in the index below, with their refresh cadence. The list updates automatically as new direct data providers are added, so it always reflects the current index.

Loading data sources…

Type: Aggregator sources aggregate and republish pricing from multiple upstream providers. Direct sources are provider-owned feeds pulled from their own endpoints or published pricing.

7.2Source Hierarchy

When a model is available from multiple sources, priority:

  1. Direct provider (e.g. openai.com pricing for GPT-5.6)
  2. Aggregator with lowest markup (e.g. OpenRouter blended price)
  3. Community submission (verified against at least one other source)

7.3Data Quality

  • Every price point stores: timestamp, source URL, raw price, normalized price
  • Automated anomaly detection: if a price moves >50% in one hour, flagged
  • Manual review of all tier additions and removals
  • Full audit trail: every index calculation is reproducible from stored raw data

8.Governance

8.1Methodology Changes

Any change to this methodology triggers:

  1. 14-day public comment period
  2. Full version increment (0.1 → 0.2)
  3. Recalculation of historical indices using new methodology
  4. Publication of both old and new values for 30-day overlap

8.2Conflict of Interest

  • InferenceIndexer is an independent verification and recommendation service
  • InferenceIndexer does not provide inference services
  • InferenceIndexer does not take positions in inference futures or derivatives
  • All data sources are public and verifiable
  • Methodology is fully transparent and reproducible

9.Limitations

9.1Known Limitations

  1. Tokenizer differences: Different models use different tokenizers. A "million tokens" from GPT-5.6 processes more text than a million tokens from Llama 3.2. This is analogous to different crude oil grades having different energy densities. The SIT accepts this imprecision as the cost of standardization.
  2. Volume data: Usage weights are sourced from OpenRouter Rankings, which covers a subset of all inference traffic. Models not available on OpenRouter are excluded from the composite basket.
  3. Aggregator dependency: Many prices are sourced via OpenRouter. If OpenRouter changes its pricing model, coverage may temporarily decrease.
  4. Quality benchmark dependency: Tier assignments depend on the Artificial Analysis Intelligence Index. Changes to their methodology affect our tiers.
  5. Excluded models: Per-request pricing, enterprise-only pricing, and deprecated models are not tracked.

9.2Future Enhancements

  • Latency-adjusted pricing (tokens/second as a factor)
  • Cache pricing tracked separately
  • Batch pricing tracked separately
  • Regional pricing (US, EU, Asia)
  • Direct provider API usage data (beyond OpenRouter)

10.Citing the SIT

When citing InferenceIndexer data in research, articles, or reports:

Text format:

InferenceIndexer SIT-Composite, August 5, 2026.
Available at: https://www.inferenceindexer.ai

Academic format:

InferenceIndexer (2026). Standard Inference Token
Methodology, v0.4.
Retrieved from https://www.inferenceindexer.ai/methodology

BibTeX:

@misc{inferenceindexer2026,
  title  = {InferenceIndexer: Standard Inference Token Methodology},
  author = {InferenceIndexer},
  year   = {2026},
  url    = {https://www.inferenceindexer.ai/methodology},
  note   = {Version 0.4}
}

11.References

  • Standard Inference Token (SIT) : Xing, Z. (23 Mar 2026) and Cunningham, M. (27 Feb 2026). EmergentMind topic summary. Defines SIT as a quality-gated inference token (MMLU >= 86%, HumanEval >= 67%, GSM8K >= 92%) and the Token Price Index (TPI) as a volume-weighted, quality-adjusted mean of spot prices. Our quality gate and adjusted price formula are adapted from this framework.
  • Artificial Analysis Intelligence Index : Independent third-party benchmark for LLM intelligence scoring. Intelligence Index v4.1 covers 257 models across 9 sub-evaluations. Used as the quality metric in our SIT formula.
  • OpenRouter API : Primary pricing data source. 400+ models from 70+ providers. Hourly price refresh. Rankings endpoint provides weekly usage data for composite weighting.

12.Harness Tables

The harness page ranks AI agent harnesses by category. It is a separate data surface from model pricing and uses a different honesty model, documented here.

Source: best-of-Agent-Harnesses, a hand-curated list of 167 harnesses rescored weekly by its maintainer (stars captured 2026-09-20 at our last refresh). We fetch its public JSON, apply the category rubrics below, and publish the result as a static table. Attribution and license (CC-BY-SA-4.0) appear on the page itself.

What is fact and what is our judgment: stars, licenses, tiers, autonomy and recovery labels are the source's facts, republished with attribution. Category fit (which harness appears under Enterprise or Small business) is OUR rubric, not a verified benchmark. The source's ranking is curation plus GitHub stars, not benchmarked evaluation; we say so on the page and do not style anything there as "verified". A research-harness table is prepared but withheld until the source list grows.

Enterprise rubric: include a harness if its license signal is open-source AND at least one of:

  • strong sandboxing (tooling_sandboxing rank 3: container/VM/OS-level isolation of agent tool execution), or
  • durable recovery (execution state survives process restarts).

Small business rubric: include a harness if ALL of:

  • it is a runnable runtime (personal-agent runtimes, coding-agent products, frameworks, multi-agent, research-task, libraries/SDKs); skill packs, curated lists, benchmark suites, observability tools, and memory layers are excluded;
  • it is managed (build-vs-buy tier 3: buy, not build) OR simple to adopt (adoption tier of mostly simple or better), AND it has at least 100 GitHub stars.

Personal agents rubric: the source's personal-agent-runtimes category verbatim, no filter: always-on, self-hosted agents run as a daemon and talked to from chat apps.

Refresh: the rubrics run weekly (Mondays) on our infrastructure and write a static snapshot; the page shows its generated-at date and the count of harnesses researched in depth by the source (87 of 167 at our last refresh). If the source changes its research coverage, the derived tables shift; the page discloses both numbers so the shift is visible, not hidden.

Contents
1.Overview2.Definition3.Quality Tiers4.Recommendation Ranking5.Index Calculation6.SIT Variants7.Data Sources8.Governance9.Limitations10.Citing the SIT11.References12.Harness Tables
InferenceIndexerInferenceIndexer.ai · Inference Recommendation Engine
ProvidersModel TypeMethodologyAPI DocsFor AgentsHarnessesData QualityAboutPrivacy PolicyTerms of Service@inferenceindex
676 models · 71 providers · Last updated: 2026-08-03 00:00 UTC