SIT Methodology
How the recommendation engine ranks models, how the Standard Inference Token price is defined, calculated, and governed.
Version 0.4, last updated August 6, 2026
1.Overview
The Standard Inference Token (SIT) tracks the marginal cost of producing AI inference tokens at a defined quality standard. The SIT Token Price Index (TPI, formerly SIT-Composite) tracks the market price of one million GPT-4-Turbo-equivalent inference tokens (1 SIT), the commodity unit for AI compute.
The SIT serves four purposes:
- Price comparison across models on a like-for-like basis
- Composite indices that track inference price movements over time
- A reference price that futures contracts can settle against
- Benchmarking: "am I paying above or below market rate?"
InferenceIndexer is an independent verification and recommendation service for AI inference. We do not provide inference services, do not route API calls, and do not take positions in any inference derivatives market. All data sources are public and verifiable.
2.Definition
2.1Unit
1 SIT = 1 million tokens of inference at a defined quality standard.
"Tokens" refers to the standard industry unit of LLM inference, as measured by each provider's tokenizer. While tokenizers differ between models, the "million tokens" convention is universally adopted and provides sufficient standardization for pricing purposes.
2.2Pricing Components
Every SIT-eligible model has three published prices:
| Component | Definition |
|---|---|
| Input price | Cost per million input (prompt) tokens |
| Output price | Cost per million output (completion) tokens |
| Blended price | Weighted average: 40% input + 60% output |
Blended price formula:
blended_price = (0.4 × input_price) + (0.6 × output_price)
The 60% output weighting reflects production workloads where output tokens exceed input tokens (generation, coding, summarization). This ratio will be refined as real-world usage data becomes available.
2.3Median Pricing (Multi-Provider Models)
Many open-source models (e.g. GLM-5.2, Llama 4, DeepSeek V4) are hosted by multiple inference providers at different prices. For these models, InferenceIndexer computes the median blended price across all available providers, rather than using a single source.
Median price formula:
model_price = median(blended_price_1, blended_price_2, ..., blended_price_n)
The median is used instead of the mean or cheapest price because it is:
- Robust to outliers: A single provider charging 10x the market rate does not skew the index.
- Not gameable: A provider cannot manipulate the headline price by temporarily dropping to $0.01.
- Realistic: The median represents what a user would realistically pay, with half the providers charging more and half charging less.
For models with only one provider, the single available price is used. The number of providers (Sources) is displayed for each model in the table and on the model detail page, where a full provider comparison table shows all available prices.
Provider endpoints are fetched daily from OpenRouter's /api/v1/models/{id}/endpoints endpoint, with direct provider APIs polled hourly. Hourly pipeline runs blend both sources and use the cached median from the most recent fetch.
3.Quality Tiers
Models are grouped into quality tiers based on demonstrated capability, using the Artificial Analysis Intelligence Index as an independent third-party benchmark. Tier boundaries are set as percentiles of the scored model population (top 10% = Frontier, next 20% = Standard, next 30% = Budget, remainder = Micro). Because the cutoffs are relative rather than fixed scores, a benchmark rebase by Artificial Analysis (e.g. Index v4.2 in September 2026) relabels no one: every model's percentile rank is unchanged by a rescale.
| Tier | Population share | Description |
|---|---|---|
| SIT-Frontier | Top 10% of scored models | Top-tier models from frontier labs |
| SIT-Standard | P70-P90 | Mid-tier production models |
| SIT-Budget | P40-P70 | Low-cost models for high-volume tasks |
| SIT-Micro | Below P40 / unscored | Ultra-cheap models for simple tasks |
Historical note (September 2026): before this change, tiers used fixed AA score thresholds (Frontier >= 50, Standard >= 30, Budget >= 15). When Artificial Analysis shipped Index v4.2 (Sep 4) and v4.3 (Sep 7) within one week, those fixed cutoffs reclassified roughly half the leaderboard twice. Percentile tiers replaced them on Sep 11, 2026.
4.Recommendation Ranking
The recommendation engine turns verified prices and quality scores into ranked, constraint-aware recommendations. It is fully deterministic: the same constraints always produce the same ranking, and no language model is involved anywhere in the ranking path. This section documents exactly how a recommendation is produced; every claim on the homepage reduces to the rules below.
4.1Eligibility
A model can appear in a recommendation only if all of the following hold. Models that fail any test are excluded before ranking, never shown with an "N/A" placeholder:
- The model is active and has a verified blended price above zero
- The model has a Cost/IQ value, which requires an Artificial Analysis Intelligence Index score (the AA ≥ 35 pipeline gate applies to having a Cost/IQ at all)
- The model is not a batch variant
- The model matches the requested modality. "text" includes vision-capable models (text+image→text), since a vision-capable model can perform text-only work
4.2Hard constraints
Callers filter the eligible pool with hard constraints. A model that fails a hard constraint is excluded, regardless of how well it would rank:
| Constraint | Meaning |
|---|---|
| budget_max_usd_per_m | Blended price per million tokens must not exceed this value |
| context_min | Context window must be at least this many tokens |
| zdr | Provider must claim zero data retention (provider-stated, not II-verified) |
| eu_sovereign | Provider must claim EU sovereignty (provider-stated, not II-verified) |
| reasoning | Include or exclude reasoning models |
| aa_min | Artificial Analysis Intelligence Index must be at least this value |
| providers | Restrict to named endpoint providers |
At least one constraint is required; unconstrained browsing is served by the model table instead. Privacy and security constraints are matched on provider statements today, not on II verification. Results say so; see the attestation table on the homepage.
4.3Ranking metric: Cost/IQ
Eligible models are ranked by Cost/IQ ascending. The formula (unchanged from v0.2, August 2026):
Cost/IQ = blended price × (40 / AA Intelligence Index score)
where the blended price is 0.4 × input price + 0.6 × output price. Lower Cost/IQ is better value: it is the verified price of a million tokens normalised per unit of demonstrated intelligence. Cost/IQ is a price-efficiency sort, not a task-fitness score. Because the formula divides by the AA score, a model with a low score and a very low price can rank first; the response carries each model's AA score so this is always visible.
4.4Task adjustments (use_case)
An optional use_case hint applies a deterministic multiplicative adjustment to Cost/IQ after the base ranking. Each profile encodes an AA floor and a reasoning preference:
| Profile | AA floor | Reasoning preference | Rationale |
|---|---|---|---|
| support | 15 | Avoid | High-volume drafting: precision and cost matter more than peak intelligence |
| volume | 10 | Avoid | High-volume generation: cost dominates |
| extraction | 20 | Avoid | Structured output: reasoning adds latency without schema accuracy gains |
| summarization | 20 | Avoid | Long context and cost efficiency matter most |
| coding | 30 | Preferred | Reasoning and higher intelligence correlate with patch quality |
| research | 35 | Preferred | Reasoning depth and large context matter most |
Penalties scale with the shortfall below the profile's AA floor; a reasoning-model penalty of 1.25 reflects that thinking tokens are not reflected in listed prices. Adjustments reorder near-ties; they never override genuine price/quality dominance. Every result shows both the raw and the adjusted basis in its "why" string.
4.5Callable-first demotion
Models with a hand-verified endpoint recipe (a tested provider base_url plus the model's native ID on that provider) rank above models without one. Demoted results carry an explicit caveat and are never silently hidden. The prefer_callable flag (default on) controls this.
4.6Freshness and honesty surface
Every recommendation carries: per-field as-of timestamps, a freshness block with the published SLA (prices 6 hours, AA scores 7 days), the runner-up models with the reason each lost, and the ranking basis in machine-readable form. Queries are never stored: the anonymous demand counter records constraint tuples only (budget, context, flags), never query content.
4.7What is NOT ranked
Latency and uptime are not scored: live probe coverage is too thin to rank on honestly (see the quality-monitoring roadmap). Security posture is not scored: no verification capability exists yet. Privacy is matched on provider statements only. When any of these changes, this section changes with them, before the homepage claims them.
5.Index Calculation
5.1Tier Indices
Each quality tier has its own index, tracking the median blended price per million tokens across all models in that tier. The SIT TPI uses equal weight per provider, capped at 30% (see Section 4.3):
| Index | What it tracks |
|---|---|
| SIT-Frontier | Median blended price of all Frontier-tier models |
| SIT-Standard | Median blended price of all Standard-tier models |
| SIT-Budget | Median blended price of all Budget-tier models |
| SIT TPI (Token Price Index) | Equal-weighted mean of cheapest SIT-qualified model per provider, 30% cap (headline price for 1 SIT) |
| SIT-Spread | Frontier price minus Budget price |
5.2Quality-Adjusted Price (Cost / IQ)
The Quality-Adjusted Price is not a transactional price. It is a normalized index value for cross-model comparison. The actual price you pay a provider is the Blended Price. The Quality-Adjusted Price normalizes that price for intelligence so models of different capability levels can be compared on a like-for-like basis.
The quality adjustment uses a transparent benchmark ratio:
- Quality gate. Only models scoring at or above the GPT-4-Turbo baseline (AA Intelligence Index >= 35) are included in the SIT TPI basket. Models below this threshold are tracked but excluded from the headline number.
- Intelligence adjustment. Prices are adjusted by the ratio of the GPT-4-Turbo reference score (40) to the model's own Artificial Analysis Intelligence Index score. A model scoring higher than GPT-4-Turbo will have a lower adjusted price (cheaper per unit of intelligence). Lower is better.
The formula:
Cost / IQ = Blended Price × (40 / AA Intelligence Score) Lower = cheaper per unit of intelligence Comparable across ALL models SIT Token Price Index (TPI) = Σ(w_p × min_adjusted_price_p) for each provider, cheapest SIT-qualified model w_p = equal weight per provider, capped at 30% eligibility: AA score in top 40% of scored models (P60+)
Where:
- Blended Price = 0.4 × input + 0.6 × output (per million tokens)
- 40 = GPT-4-Turbo (Jan 2024) reference score on AA Intelligence Index v4.1. Models scoring 40 are at GPT-4-Turbo parity. The reference is based on the benchmark thresholds defined in the SIT standard (MMLU >= 86%, HumanEval >= 67%, GSM8K >= 92%).
- AA Intelligence Index = Artificial Analysis Intelligence Index v4.1, an independent third-party benchmark
Lower Cost / IQ = cheaper per unit of intelligence. Cost / IQ is an absolute measure, not relative to any tier median. A model at $0.25/M Cost / IQ is cheaper per unit of intelligence than a model at $1.50/M Cost / IQ, regardless of tier. Models without an AA Intelligence Index score do not receive a Cost / IQ and are excluded from the composite basket.
5.3Provider Equal Weighting (TPI)
The SIT TPI uses a provider equal weighting: each provider contributes their cheapest SIT-qualified model to the basket, and every provider carries equal weight (capped at 30% of total weight). This is the Token Price Index (TPI) methodology from the SIT paper (arXiv:2603.21690).
Equal weight per provider prevents any single provider from dominating the index regardless of how many models they host. The 30% cap prevents large providers from overwhelming the index as more providers are added.
For each provider: cheapest SIT-qualified model's Cost / IQ TPI = Σ(w_p × min_adjusted_price_p) w_p = equal weight per provider, capped at 30% eligibility: AA score in top 40% of scored models (P60+)
Per-tier indices (Frontier, Standard, Budget, Micro) use a simple median across all models in that tier. This answers: "What does a typical model in this tier cost?" The SIT TPI uses provider equal weighting to answer: "What does the market charge for GPT-4-equivalent inference?"
5.4Calculation Frequency
| Frequency | What happens |
|---|---|
| Hourly | Pull pricing from all sources, update database |
| Daily | Calculate SIT indices, publish at 00:00 UTC |
| Weekly | Refresh usage weights from OpenRouter Rankings (Mondays 06:00 UTC) |
| Monthly | Review tier composition, add/remove models |
5.5Base Date and Rebaselining
- Current base date: September 4, 2026 (SIT TPI = 1000 at this date). Charts mark era breaks where the underlying Artificial Analysis Index version changed.
- Prior base: August 4, 2026 (superseded Sep 11, 2026).
- Rebaselining on methodology changes. All rebaselining events are published with full explanation. The September 2026 rebase was forced by two Artificial Analysis Index updates in one week (v4.2 on Sep 4, v4.3 on Sep 7), which changed the composition of the index basket without any real prices moving.
6.SIT Variants
Inference is not a single homogeneous commodity. The SIT supports attribute-based filtering, similar to how CoinMarketCap filters by category (DeFi, Layer 1, etc.).
| SIT Variant | Filter |
|---|---|
| SIT-Composite | All models (headline number) |
| SIT-Frontier | AA Index >= 50 |
| SIT-Standard | AA Index 30–49 |
| SIT-Budget | AA Index 15–29 |
| SIT-EU-Sovereign | EU-hosted only (jurisdiction) |
| SIT-ZDR | Zero data retention guaranteed (provider does not store or train on inputs) |
| SIT-Open | Open weights models only |
| SIT-Proprietary | Proprietary models only |
| SIT-Cached | With prompt caching applied (7:2:1 blend) |
The SIT TPI is always the headline number. Variant indices allow users to track specific segments of the inference market.
ZDR and EU Infra status: Provider classifications for SIT-ZDR and SIT-EU-Sovereign are based on publicly available provider documentation as of August 2026. ZDR (Zero Data Retention) includes providers that do not store or train on user inputs by default on their standard API. EU-Sovereign includes providers domiciled in the EU/EEA that are not subject to the US CLOUD Act. Many providers offer ZDR or EU hosting only on enterprise plans; these are not classified as ZDR or EU-Sovereign here. Provider policies change; classifications are reviewed quarterly.
7.Data Sources
7.1Primary Sources
All providers in the index below, with their refresh cadence. The list updates automatically as new direct data providers are added, so it always reflects the current index.
Loading data sources…
Type: Aggregator sources aggregate and republish pricing from multiple upstream providers. Direct sources are provider-owned feeds pulled from their own endpoints or published pricing.
7.2Source Hierarchy
When a model is available from multiple sources, priority:
- Direct provider (e.g. openai.com pricing for GPT-5.6)
- Aggregator with lowest markup (e.g. OpenRouter blended price)
- Community submission (verified against at least one other source)
7.3Data Quality
- Every price point stores: timestamp, source URL, raw price, normalized price
- Automated anomaly detection: if a price moves >50% in one hour, flagged
- Manual review of all tier additions and removals
- Full audit trail: every index calculation is reproducible from stored raw data
8.Governance
8.1Methodology Changes
Any change to this methodology triggers:
- 14-day public comment period
- Full version increment (0.1 → 0.2)
- Recalculation of historical indices using new methodology
- Publication of both old and new values for 30-day overlap
8.2Conflict of Interest
- InferenceIndexer is an independent verification and recommendation service
- InferenceIndexer does not provide inference services
- InferenceIndexer does not take positions in inference futures or derivatives
- All data sources are public and verifiable
- Methodology is fully transparent and reproducible
9.Limitations
9.1Known Limitations
- Tokenizer differences: Different models use different tokenizers. A "million tokens" from GPT-5.6 processes more text than a million tokens from Llama 3.2. This is analogous to different crude oil grades having different energy densities. The SIT accepts this imprecision as the cost of standardization.
- Volume data: Usage weights are sourced from OpenRouter Rankings, which covers a subset of all inference traffic. Models not available on OpenRouter are excluded from the composite basket.
- Aggregator dependency: Many prices are sourced via OpenRouter. If OpenRouter changes its pricing model, coverage may temporarily decrease.
- Quality benchmark dependency: Tier assignments depend on the Artificial Analysis Intelligence Index. Changes to their methodology affect our tiers.
- Excluded models: Per-request pricing, enterprise-only pricing, and deprecated models are not tracked.
9.2Future Enhancements
- Latency-adjusted pricing (tokens/second as a factor)
- Cache pricing tracked separately
- Batch pricing tracked separately
- Regional pricing (US, EU, Asia)
- Direct provider API usage data (beyond OpenRouter)
10.Citing the SIT
When citing InferenceIndexer data in research, articles, or reports:
Text format:
InferenceIndexer SIT-Composite, August 5, 2026. Available at: https://www.inferenceindexer.ai
Academic format:
InferenceIndexer (2026). Standard Inference Token Methodology, v0.4. Retrieved from https://www.inferenceindexer.ai/methodology
BibTeX:
@misc{inferenceindexer2026,
title = {InferenceIndexer: Standard Inference Token Methodology},
author = {InferenceIndexer},
year = {2026},
url = {https://www.inferenceindexer.ai/methodology},
note = {Version 0.4}
}11.References
- Standard Inference Token (SIT) : Xing, Z. (23 Mar 2026) and Cunningham, M. (27 Feb 2026). EmergentMind topic summary. Defines SIT as a quality-gated inference token (MMLU >= 86%, HumanEval >= 67%, GSM8K >= 92%) and the Token Price Index (TPI) as a volume-weighted, quality-adjusted mean of spot prices. Our quality gate and adjusted price formula are adapted from this framework.
- Artificial Analysis Intelligence Index : Independent third-party benchmark for LLM intelligence scoring. Intelligence Index v4.1 covers 257 models across 9 sub-evaluations. Used as the quality metric in our SIT formula.
- OpenRouter API : Primary pricing data source. 400+ models from 70+ providers. Hourly price refresh. Rankings endpoint provides weekly usage data for composite weighting.
12.Harness Tables
The harness page ranks AI agent harnesses by category. It is a separate data surface from model pricing and uses a different honesty model, documented here.
Source: best-of-Agent-Harnesses, a hand-curated list of 167 harnesses rescored weekly by its maintainer (stars captured 2026-09-20 at our last refresh). We fetch its public JSON, apply the category rubrics below, and publish the result as a static table. Attribution and license (CC-BY-SA-4.0) appear on the page itself.
What is fact and what is our judgment: stars, licenses, tiers, autonomy and recovery labels are the source's facts, republished with attribution. Category fit (which harness appears under Enterprise or Small business) is OUR rubric, not a verified benchmark. The source's ranking is curation plus GitHub stars, not benchmarked evaluation; we say so on the page and do not style anything there as "verified". A research-harness table is prepared but withheld until the source list grows.
Enterprise rubric: include a harness if its license signal is open-source AND at least one of:
- strong sandboxing (tooling_sandboxing rank 3: container/VM/OS-level isolation of agent tool execution), or
- durable recovery (execution state survives process restarts).
Small business rubric: include a harness if ALL of:
- it is a runnable runtime (personal-agent runtimes, coding-agent products, frameworks, multi-agent, research-task, libraries/SDKs); skill packs, curated lists, benchmark suites, observability tools, and memory layers are excluded;
- it is managed (build-vs-buy tier 3: buy, not build) OR simple to adopt (adoption tier of mostly simple or better), AND it has at least 100 GitHub stars.
Personal agents rubric: the source's personal-agent-runtimes category verbatim, no filter: always-on, self-hosted agents run as a daemon and talked to from chat apps.
Refresh: the rubrics run weekly (Mondays) on our infrastructure and write a static snapshot; the page shows its generated-at date and the count of harnesses researched in depth by the source (87 of 167 at our last refresh). If the source changes its research coverage, the derived tables shift; the page discloses both numbers so the shift is visible, not hidden.