InferenceIndexerInferenceIndexer.ai
/
ModelsProvidersAPIEmbedFor agentsHarnessesMethodologyAboutLoginSign Up

API Documentation

Recommendations, pricing, and verification data. 1,000 requests/day with a free API key.

Get API Key
Authentication
Getting a key
Endpoints
/recommend/explain/sit/composite/latest/sit/composite/history/models/models/{id}/models/{id}/history
Reference
Rate LimitsResponse FormatErrors

Authentication

API keys are required for all requests. For agents, get one instantly with no email or account:

$ curl -X POST https://api.inferenceindexer.ai/v1/auth/anonymous

Returns a key valid on the free tier (10,000 requests/day, 30 days of history).

Or sign up with an email (usage tracking + dashboard):

  1. 1Click "Get API Key" or go to the Sign Up page
  2. 2Enter your email address
  3. 3Click the magic link in the email
  4. 4Your API key is displayed on the dashboard

Pass your API key in the Authorization header:

Authorization: Bearer YOUR_API_KEY

Example:

$ curl -H "Authorization: Bearer YOUR_API_KEY" \
     https://api.inferenceindexer.ai/v1/sit/composite/latest

Endpoints

Each endpoint is a card. All paths are prefixed with the API base URL.

GET/v1/sit/composite/latest

Returns the current SIT Token Price Index (TPI), including tier breakdowns. This is the market price for GPT-4-equivalent inference (1 SIT).

Parameters

None

Example Request

$ curl -H "Authorization: Bearer YOUR_API_KEY" \
     https://api.inferenceindexer.ai/v1/sit/composite/latest

Example Response (200 OK)

{
  "date": "2026-08-03",
  "composite": {
    "price_per_m": 2.84,
    "index_points": 784.5,
    "change_24h": -1.2,
    "change_7d": -3.1,
    "change_30d": -12.4
  },
  "tiers": {
    "frontier": { "price_per_m": 35.20, "change_24h": -0.8, "models": 8 },
    "standard": { "price_per_m": 1.25, "change_24h": -1.5, "models": 156 },
    "budget": { "price_per_m": 0.42, "change_24h": -2.1, "models": 78 }
  },
  "spread": { "price_per_m": 34.78, "change_24h": -1.9 }
}
POST/v1/recommend

Constraint-filtered model ranking with receipts. Ranks by Cost/IQ (blended price x 40/AA score, lower is better). Returns ranked models with a plain-English 'why', a hot-swap endpoint_config (provider base_url + native model id), per-field as-of timestamps, and runner-ups with reasons. Requires at least one constraint.

Parameters

None

Example Request

$ curl -X POST -H "Authorization: Bearer ***" -H "Content-Type: application/json" \\
     -d '{"budget_max_usd_per_m": 1, "context_min": 100000, "limit": 3}' \\
     "https://api.inferenceindexer.ai/v1/recommend"

Example Response (200 OK)

{
  "query": { "budget_max_usd_per_m": 1, "context_min": 100000, "modality": "text" },
  "methodology_version": "0.2",
  "recommendations": [
    {
      "rank": 1,
      "model_id": "deepseek/deepseek-v4-flash",
      "tier": "frontier",
      "blended_price_per_m": 0.224,
      "cost_per_iq": 0.173,
      "context_length": 1048576,
      "is_reasoning": false,
      "why": "Rank 1 by Cost/IQ ($0.173/GPT-4-equiv M tokens) of models under $1.0/M blended and >= 100,000 token context; cheapest verified endpoint at DeepInfra ($0.224/M blended)",
      "endpoint_config": {
        "provider": "DeepInfra",
        "base_url": "https://api.deepinfra.com/v1/openai",
        "model_id_on_provider": "deepseek-ai/DeepSeek-V4-Flash",
        "model_id_source": "verified",
        "input_price_per_m": 0.07,
        "output_price_per_m": 0.4,
        "auth_scheme": "bearer"
      },
      "as_of": { "price": "2026-09-02T16:00:25Z", "aa_score": "2026-09-01" },
      "caveats": []
    }
  ],
  "alternatives_considered": { "count": 45, "runner_ups": [ ... ] },
  "ranking_basis": {
    "method": "cost_per_iq_ascending",
    "description": "Blended price x (40 / AA Intelligence score). Quality gate AA >= 35.",
    "quality_latency_data": "Not used in v1 ranking: probe coverage too thin."
  },
  "disclaimer": "Estimates based on aggregated public pricing. Verify with the provider before committing spend."
}
GET/v1/explain?model_id={id}

One call answers 'what is this model like right now': current pricing with 24h/7d changes, price history summary with trend, all endpoints across providers, the cheapest hand-verified endpoint (with native model id for hot-swapping), privacy flags (ZDR/EU), and the AA intelligence score. Everything as-of stamped.

Parameters

ParameterTypeRequiredDescription
model_idstringYesCanonical model id, e.g. anthropic/claude-sonnet-5
history_daysintegerNoHistory window (default 30, max 365)

Example Request

$ curl -H "Authorization: Bearer ***" \\
     "https://api.inferenceindexer.ai/v1/explain?model_id=deepseek/deepseek-v4-flash"

Example Response (200 OK)

{
  "model_id": "deepseek/deepseek-v4-flash",
  "tier": "frontier",
  "as_of": "2026-09-02T16:00:25Z",
  "pricing": {
    "input_per_m": 0.1, "output_per_m": 0.4,
    "blended_per_m": 0.224, "cost_per_iq": 0.173,
    "change_24h": 0.0, "change_7d": 0.0
  },
  "price_history_summary": {
    "days_observed": 30, "snapshots": 763,
    "min_blended": 0.2202, "max_blended": 0.224,
    "trend": "flat", "trend_note": "estimate: first-3 vs last-3 average in window",
    "series_url": "/v1/models/deepseek/deepseek-v4-flash/history?days=30"
  },
  "endpoints": [ { "provider": "DeepInfra", "blended_per_m": 0.224 }, ... ],
  "cheapest_verified_endpoint": {
    "provider": "DeepInfra",
    "base_url": "https://api.deepinfra.com/v1/openai",
    "model_id_on_provider": "deepseek-ai/DeepSeek-V4-Flash",
    "model_id_source": "verified",
    "blended_per_m": 0.224,
    "auth_scheme": "bearer"
  },
  "privacy": { "zdr_available": true, "eu_sovereign_available": false, "note": "Provider-level claims, aggregated not certified." },
  "quality": { "aa_score": 51.77, "cost_per_iq": 0.173, "probe_data": "insufficient coverage" },
  "methodology_version": "0.2"
}
GET/v1/sit/composite/history?days=30

Returns historical SIT TPI values. Alias endpoint: GET /v1/tpi/history.

Parameters

ParameterTypeRequiredDescription
daysintegerNoNumber of days to return (default: 30, max: 365)
tierstringNoFilter to a specific tier: frontier, standard, budget

Example Request

$ curl -H "Authorization: Bearer YOUR_API_KEY" \
     "https://api.inferenceindexer.ai/v1/sit/composite/history?days=30"

Example Response (200 OK)

{
  "history": [
    { "date": "2026-08-03", "price_per_m": 2.84, "index_points": 784.5 },
    { "date": "2026-08-02", "price_per_m": 2.87, "index_points": 792.8 },
    ...
  ]
}
GET/v1/models?tier=standard&sort=blended

Returns all tracked models with current pricing.

Parameters

ParameterTypeRequiredDescription
tierstringNoFilter by tier: frontier, standard, budget, micro
providerstringNoFilter by provider name
sortstringNoSort by: blended, input, output, sit_adjusted_price (default: sit_adjusted_price)
limitintegerNoMax results (default: 50, max: 500)

Example Request

$ curl -H "Authorization: Bearer YOUR_API_KEY" \
     "https://api.inferenceindexer.ai/v1/models?tier=standard&sort=blended"

Example Response (200 OK)

{
  "count": 318,
  "models": [
    {
      "model_id": "deepseek/deepseek-v4-reasoner",
      "name": "DeepSeek V4 Reasoner",
      "provider": "DeepSeek",
      "tier": "frontier",
      "input_price_per_m": 0.55,
      "output_price_per_m": 2.19,
      "blended_price_per_m": 1.53,
      "sit_adjusted_price": 0.04,
      "context_length": 128000,
      "change_24h": -3.0,
      "change_7d": -8.0
    },
    ...
  ]
}
GET/v1/models/{model_id}

Returns detailed pricing and metadata for a single model.

Parameters

ParameterTypeRequiredDescription
model_idstringYesThe model ID (e.g. openai/gpt-5.6)

Example Request

$ curl -H "Authorization: Bearer YOUR_API_KEY" \
     https://api.inferenceindexer.ai/v1/models/openai/gpt-5.6

Example Response (200 OK)

{
  "model_id": "openai/gpt-5.6",
  "name": "OpenAI: GPT-5.6",
  "provider": "OpenAI",
  "tier": "frontier",
  "input_price_per_m": 2.50,
  "output_price_per_m": 10.00,
  "blended_price_per_m": 4.75,
  "sit_score": 0.12,
  "context_length": 256000,
  "change_24h": 0.0,
  "change_7d": -2.1
}
GET/v1/models/{model_id}/history?days=90

Returns historical price data for a single model.

Parameters

ParameterTypeRequiredDescription
model_idstringYesThe model ID
daysintegerNoDays of history (default: 30, max: 365 on free tier)

Example Request

$ curl -H "Authorization: Bearer YOUR_API_KEY" \
     "https://api.inferenceindexer.ai/v1/models/openai/gpt-5.6/history?days=90"

Example Response (200 OK)

{
  "model_id": "openai/gpt-5.6",
  "history": [
    { "date": "2026-08-03", "blended_price_per_m": 4.75 },
    { "date": "2026-08-02", "blended_price_per_m": 4.80 },
    ...
  ]
}

Rate Limits

PlanRequests/dayRequests/minuteHistory access
Public (no key)100107 days
Free (email)1,0003030 days
Paid (future)50,000100365 days

Rate limit headers are included in every response:

X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 987
X-RateLimit-Reset: 1691078400

Response Format

All responses are JSON. Timestamps are ISO 8601 UTC. Prices are in USD per million tokens.

Successful response (200 OK)

{
  "data": { ... },
  "meta": {
    "request_id": "req_abc123",
    "timestamp": "2026-08-03T15:00:00Z"
  }
}

Pagination (for list endpoints)

{
  "count": 318,
  "page": 1,
  "per_page": 50,
  "total_pages": 7,
  "next": "/v1/models?page=2"
}

Errors

Status CodeMeaning
200 OKSuccess
400 Bad RequestInvalid parameter (check the error message)
401 UnauthorizedMissing or invalid API key
404 Not FoundModel or endpoint doesn't exist
429 Too ManyRate limit exceeded. Check X-RateLimit-Reset header
500 Server ErrorSomething went wrong on our end. Retry after a few seconds

Error response format

{
  "error": {
    "code": "rate_limit_exceeded",
    "message": "Rate limit of 1000 requests/day exceeded. Resets at 2026-08-04T00:00:00Z.",
    "documentation_url": "https://www.inferenceindexer.ai/api/docs#rate-limits"
  }
}