Authentication
API keys are required for all requests. For agents, get one instantly with no email or account:
$ curl -X POST https://api.inferenceindexer.ai/v1/auth/anonymousReturns a key valid on the free tier (10,000 requests/day, 30 days of history).
Or sign up with an email (usage tracking + dashboard):
- 1Click "Get API Key" or go to the Sign Up page
- 2Enter your email address
- 3Click the magic link in the email
- 4Your API key is displayed on the dashboard
Pass your API key in the Authorization header:
Authorization: Bearer YOUR_API_KEY
Example:
$ curl -H "Authorization: Bearer YOUR_API_KEY" \
https://api.inferenceindexer.ai/v1/sit/composite/latestEndpoints
Each endpoint is a card. All paths are prefixed with the API base URL.
/v1/sit/composite/latestReturns the current SIT Token Price Index (TPI), including tier breakdowns. This is the market price for GPT-4-equivalent inference (1 SIT).
Parameters
None
Example Request
$ curl -H "Authorization: Bearer YOUR_API_KEY" \
https://api.inferenceindexer.ai/v1/sit/composite/latestExample Response (200 OK)
{
"date": "2026-08-03",
"composite": {
"price_per_m": 2.84,
"index_points": 784.5,
"change_24h": -1.2,
"change_7d": -3.1,
"change_30d": -12.4
},
"tiers": {
"frontier": { "price_per_m": 35.20, "change_24h": -0.8, "models": 8 },
"standard": { "price_per_m": 1.25, "change_24h": -1.5, "models": 156 },
"budget": { "price_per_m": 0.42, "change_24h": -2.1, "models": 78 }
},
"spread": { "price_per_m": 34.78, "change_24h": -1.9 }
}/v1/recommendConstraint-filtered model ranking with receipts. Ranks by Cost/IQ (blended price x 40/AA score, lower is better). Returns ranked models with a plain-English 'why', a hot-swap endpoint_config (provider base_url + native model id), per-field as-of timestamps, and runner-ups with reasons. Requires at least one constraint.
Parameters
None
Example Request
$ curl -X POST -H "Authorization: Bearer ***" -H "Content-Type: application/json" \\
-d '{"budget_max_usd_per_m": 1, "context_min": 100000, "limit": 3}' \\
"https://api.inferenceindexer.ai/v1/recommend"Example Response (200 OK)
{
"query": { "budget_max_usd_per_m": 1, "context_min": 100000, "modality": "text" },
"methodology_version": "0.2",
"recommendations": [
{
"rank": 1,
"model_id": "deepseek/deepseek-v4-flash",
"tier": "frontier",
"blended_price_per_m": 0.224,
"cost_per_iq": 0.173,
"context_length": 1048576,
"is_reasoning": false,
"why": "Rank 1 by Cost/IQ ($0.173/GPT-4-equiv M tokens) of models under $1.0/M blended and >= 100,000 token context; cheapest verified endpoint at DeepInfra ($0.224/M blended)",
"endpoint_config": {
"provider": "DeepInfra",
"base_url": "https://api.deepinfra.com/v1/openai",
"model_id_on_provider": "deepseek-ai/DeepSeek-V4-Flash",
"model_id_source": "verified",
"input_price_per_m": 0.07,
"output_price_per_m": 0.4,
"auth_scheme": "bearer"
},
"as_of": { "price": "2026-09-02T16:00:25Z", "aa_score": "2026-09-01" },
"caveats": []
}
],
"alternatives_considered": { "count": 45, "runner_ups": [ ... ] },
"ranking_basis": {
"method": "cost_per_iq_ascending",
"description": "Blended price x (40 / AA Intelligence score). Quality gate AA >= 35.",
"quality_latency_data": "Not used in v1 ranking: probe coverage too thin."
},
"disclaimer": "Estimates based on aggregated public pricing. Verify with the provider before committing spend."
}/v1/explain?model_id={id}One call answers 'what is this model like right now': current pricing with 24h/7d changes, price history summary with trend, all endpoints across providers, the cheapest hand-verified endpoint (with native model id for hot-swapping), privacy flags (ZDR/EU), and the AA intelligence score. Everything as-of stamped.
Parameters
Example Request
$ curl -H "Authorization: Bearer ***" \\
"https://api.inferenceindexer.ai/v1/explain?model_id=deepseek/deepseek-v4-flash"Example Response (200 OK)
{
"model_id": "deepseek/deepseek-v4-flash",
"tier": "frontier",
"as_of": "2026-09-02T16:00:25Z",
"pricing": {
"input_per_m": 0.1, "output_per_m": 0.4,
"blended_per_m": 0.224, "cost_per_iq": 0.173,
"change_24h": 0.0, "change_7d": 0.0
},
"price_history_summary": {
"days_observed": 30, "snapshots": 763,
"min_blended": 0.2202, "max_blended": 0.224,
"trend": "flat", "trend_note": "estimate: first-3 vs last-3 average in window",
"series_url": "/v1/models/deepseek/deepseek-v4-flash/history?days=30"
},
"endpoints": [ { "provider": "DeepInfra", "blended_per_m": 0.224 }, ... ],
"cheapest_verified_endpoint": {
"provider": "DeepInfra",
"base_url": "https://api.deepinfra.com/v1/openai",
"model_id_on_provider": "deepseek-ai/DeepSeek-V4-Flash",
"model_id_source": "verified",
"blended_per_m": 0.224,
"auth_scheme": "bearer"
},
"privacy": { "zdr_available": true, "eu_sovereign_available": false, "note": "Provider-level claims, aggregated not certified." },
"quality": { "aa_score": 51.77, "cost_per_iq": 0.173, "probe_data": "insufficient coverage" },
"methodology_version": "0.2"
}/v1/sit/composite/history?days=30Returns historical SIT TPI values. Alias endpoint: GET /v1/tpi/history.
Parameters
Example Request
$ curl -H "Authorization: Bearer YOUR_API_KEY" \
"https://api.inferenceindexer.ai/v1/sit/composite/history?days=30"Example Response (200 OK)
{
"history": [
{ "date": "2026-08-03", "price_per_m": 2.84, "index_points": 784.5 },
{ "date": "2026-08-02", "price_per_m": 2.87, "index_points": 792.8 },
...
]
}/v1/models?tier=standard&sort=blendedReturns all tracked models with current pricing.
Parameters
Example Request
$ curl -H "Authorization: Bearer YOUR_API_KEY" \
"https://api.inferenceindexer.ai/v1/models?tier=standard&sort=blended"Example Response (200 OK)
{
"count": 318,
"models": [
{
"model_id": "deepseek/deepseek-v4-reasoner",
"name": "DeepSeek V4 Reasoner",
"provider": "DeepSeek",
"tier": "frontier",
"input_price_per_m": 0.55,
"output_price_per_m": 2.19,
"blended_price_per_m": 1.53,
"sit_adjusted_price": 0.04,
"context_length": 128000,
"change_24h": -3.0,
"change_7d": -8.0
},
...
]
}/v1/models/{model_id}Returns detailed pricing and metadata for a single model.
Parameters
Example Request
$ curl -H "Authorization: Bearer YOUR_API_KEY" \
https://api.inferenceindexer.ai/v1/models/openai/gpt-5.6Example Response (200 OK)
{
"model_id": "openai/gpt-5.6",
"name": "OpenAI: GPT-5.6",
"provider": "OpenAI",
"tier": "frontier",
"input_price_per_m": 2.50,
"output_price_per_m": 10.00,
"blended_price_per_m": 4.75,
"sit_score": 0.12,
"context_length": 256000,
"change_24h": 0.0,
"change_7d": -2.1
}/v1/models/{model_id}/history?days=90Returns historical price data for a single model.
Parameters
Example Request
$ curl -H "Authorization: Bearer YOUR_API_KEY" \
"https://api.inferenceindexer.ai/v1/models/openai/gpt-5.6/history?days=90"Example Response (200 OK)
{
"model_id": "openai/gpt-5.6",
"history": [
{ "date": "2026-08-03", "blended_price_per_m": 4.75 },
{ "date": "2026-08-02", "blended_price_per_m": 4.80 },
...
]
}Rate Limits
Rate limit headers are included in every response:
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 987
X-RateLimit-Reset: 1691078400Response Format
All responses are JSON. Timestamps are ISO 8601 UTC. Prices are in USD per million tokens.
Successful response (200 OK)
{
"data": { ... },
"meta": {
"request_id": "req_abc123",
"timestamp": "2026-08-03T15:00:00Z"
}
}Pagination (for list endpoints)
{
"count": 318,
"page": 1,
"per_page": 50,
"total_pages": 7,
"next": "/v1/models?page=2"
}Errors
Error response format
{
"error": {
"code": "rate_limit_exceeded",
"message": "Rate limit of 1000 requests/day exceeded. Resets at 2026-08-04T00:00:00Z.",
"documentation_url": "https://www.inferenceindexer.ai/api/docs#rate-limits"
}
}