Pricing & billing
Prepaid credits, per-1M-token rates, separate input/output/cached pricing, balance and spend limits.
Pricing & billing
Inference is prepaid. You top up a credit balance in the Dashboard, and each request deducts from that balance based on actual token usage. No subscription, no minimum spend.
The unit of pricing
All prices are expressed in US dollars per 1,000,000 tokens (1M tokens), split three ways:
| Category | Description |
|---|---|
| Input | Prompt tokens you send, including system messages, conversation history, tool definitions and tokens derived from images. |
| Cached input | Input tokens served from the upstream prompt cache, priced substantially below regular input. |
| Output | Tokens the model generates — usually the most expensive of the three. |
Cost for a single request:
cost = (input_tokens / 1_000_000) × input_per_1m
+ (cached_input_tokens / 1_000_000) × cached_input_per_1m
+ (output_tokens / 1_000_000) × output_per_1m
Platform markup
We apply a fixed-percentage platform markup on top of the upstream provider's cost. It funds the gateway's operations, failover and billing. The pricing returned by GET /v1/models already includes the markup — it is exactly the rate you are charged, with no additional hidden fees.
The markup rate is uniform across models and is published on the Pricing page in the Dashboard.
Example rates
The figures below are illustrative. The authoritative rates always come from GET /v1/models and the Dashboard.
| Model | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
openai/gpt-4o | $2.50 | $1.25 | $10.00 |
openai/gpt-4o-mini | $0.15 | $0.075 | $0.60 |
anthropic/claude-sonnet-4-5 | $3.00 | $0.30 | $15.00 |
google/gemini-2.5-pro | $1.25 | $0.31 | $10.00 |
google/gemini-2.5-flash | $0.30 | $0.075 | $2.50 |
Computing cost from a response
The usage block on every response is what we bill on:
{
"usage": {
"prompt_tokens": 12000,
"completion_tokens": 800,
"total_tokens": 12800,
"prompt_tokens_details": { "cached_tokens": 9000 }
}
}
Note that prompt_tokens includes cached_tokens, so subtract before pricing:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
PRICING = { # USD per 1M tokens; in production, read this from /v1/models
"openai/gpt-4o": {"input": 2.50, "cached_input": 1.25, "output": 10.00},
}
resp = client.chat.completions.create(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "Hello"}],
)
u = resp.usage
cached = (u.prompt_tokens_details.cached_tokens if u.prompt_tokens_details else 0) or 0
fresh_input = u.prompt_tokens - cached
p = PRICING[resp.model]
cost = (
fresh_input / 1_000_000 * p["input"]
+ cached / 1_000_000 * p["cached_input"]
+ u.completion_tokens / 1_000_000 * p["output"]
)
print(f"cost: ${cost:.6f}")
The same calculation in Node:
const u = resp.usage!;
const cached = u.prompt_tokens_details?.cached_tokens ?? 0;
const freshInput = u.prompt_tokens - cached;
const cost =
(freshInput / 1_000_000) * 2.5 +
(cached / 1_000_000) * 1.25 +
(u.completion_tokens / 1_000_000) * 10.0;
console.log(`cost: $${cost.toFixed(6)}`);
Insufficient balance
If your balance cannot cover a request, it is rejected outright:
HTTP/1.1 402 Payment Required
Content-Type: application/json
{
"error": {
"code": "insufficient_credit",
"message": "Your credit balance is insufficient for this request."
}
}
Top up under Billing in the Dashboard to resume. Enable low-balance notifications so production never stops unexpectedly.
Monthly spend limits
Every key can carry its own monthly spend limit in USD. Once the month's cost reaches it:
HTTP/1.1 402 Payment Required
Content-Type: application/json
{
"error": {
"code": "spend_limit_reached",
"message": "This API key has reached its monthly spend limit of $50.00."
}
}
Limits reset automatically at the start of each month. This is the single most effective guard against a runaway loop or a leaked key — set one on every key that leaves your machine.
Usage reporting
The Usage page in the Dashboard provides:
- Requests, tokens and cost grouped by date, key and model.
- Per-request detail: timestamp, model, input / cached / output tokens, and the cost of that request.
- CSV export for internal chargeback.
Cost-saving techniques
| Technique | Effect |
|---|---|
| Match model to task difficulty | Classification, extraction and reformatting on a small model often costs under a tenth of a flagship. |
| Keep prompt prefixes stable | Put system messages and few-shot examples first and leave them unchanged to hit cached input. |
Set max_tokens | Output is the priciest category; a cap prevents runaway responses. |
| Trim conversation history | Summarise long threads instead of accumulating messages indefinitely. |
| Keep the tool list short | Every tool definition consumes input tokens. |
| Downscale images | Images convert to input tokens by resolution; oversized images are pure waste. |
Common questions
- Are failed requests billed? Requests rejected before reaching the upstream — 401, 402, 403, 404, 429 — are not billed. If generation had already begun before failure (
upstream_error), the tokens produced are billed. - Is an aborted stream billed? Yes. Tokens the upstream already generated are billed.
- Do credits expire? No.