Skip to content

Pricing & billing

Prepaid credits, per-1M-token rates, separate input/output/cached pricing, balance and spend limits.

Pricing & billing

Inference is prepaid. You top up a credit balance in the Dashboard, and each request deducts from that balance based on actual token usage. No subscription, no minimum spend.

The unit of pricing

All prices are expressed in US dollars per 1,000,000 tokens (1M tokens), split three ways:

CategoryDescription
InputPrompt tokens you send, including system messages, conversation history, tool definitions and tokens derived from images.
Cached inputInput tokens served from the upstream prompt cache, priced substantially below regular input.
OutputTokens the model generates — usually the most expensive of the three.

Cost for a single request:

cost = (input_tokens        / 1_000_000) × input_per_1m
     + (cached_input_tokens / 1_000_000) × cached_input_per_1m
     + (output_tokens       / 1_000_000) × output_per_1m

Platform markup

We apply a fixed-percentage platform markup on top of the upstream provider's cost. It funds the gateway's operations, failover and billing. The pricing returned by GET /v1/models already includes the markup — it is exactly the rate you are charged, with no additional hidden fees.

The markup rate is uniform across models and is published on the Pricing page in the Dashboard.

Example rates

The figures below are illustrative. The authoritative rates always come from GET /v1/models and the Dashboard.

ModelInput / 1MCached input / 1MOutput / 1M
openai/gpt-4o$2.50$1.25$10.00
openai/gpt-4o-mini$0.15$0.075$0.60
anthropic/claude-sonnet-4-5$3.00$0.30$15.00
google/gemini-2.5-pro$1.25$0.31$10.00
google/gemini-2.5-flash$0.30$0.075$2.50

Computing cost from a response

The usage block on every response is what we bill on:

{
  "usage": {
    "prompt_tokens": 12000,
    "completion_tokens": 800,
    "total_tokens": 12800,
    "prompt_tokens_details": { "cached_tokens": 9000 }
  }
}

Note that prompt_tokens includes cached_tokens, so subtract before pricing:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.alphacurve.io/v1",
    api_key=os.environ["INFERENCE_API_KEY"],
)

PRICING = {  # USD per 1M tokens; in production, read this from /v1/models
    "openai/gpt-4o": {"input": 2.50, "cached_input": 1.25, "output": 10.00},
}

resp = client.chat.completions.create(
    model="openai/gpt-4o",
    messages=[{"role": "user", "content": "Hello"}],
)

u = resp.usage
cached = (u.prompt_tokens_details.cached_tokens if u.prompt_tokens_details else 0) or 0
fresh_input = u.prompt_tokens - cached
p = PRICING[resp.model]

cost = (
    fresh_input / 1_000_000 * p["input"]
    + cached / 1_000_000 * p["cached_input"]
    + u.completion_tokens / 1_000_000 * p["output"]
)
print(f"cost: ${cost:.6f}")

The same calculation in Node:

const u = resp.usage!;
const cached = u.prompt_tokens_details?.cached_tokens ?? 0;
const freshInput = u.prompt_tokens - cached;

const cost =
  (freshInput / 1_000_000) * 2.5 +
  (cached / 1_000_000) * 1.25 +
  (u.completion_tokens / 1_000_000) * 10.0;

console.log(`cost: $${cost.toFixed(6)}`);

Insufficient balance

If your balance cannot cover a request, it is rejected outright:

HTTP/1.1 402 Payment Required
Content-Type: application/json

{
  "error": {
    "code": "insufficient_credit",
    "message": "Your credit balance is insufficient for this request."
  }
}

Top up under Billing in the Dashboard to resume. Enable low-balance notifications so production never stops unexpectedly.

Monthly spend limits

Every key can carry its own monthly spend limit in USD. Once the month's cost reaches it:

HTTP/1.1 402 Payment Required
Content-Type: application/json

{
  "error": {
    "code": "spend_limit_reached",
    "message": "This API key has reached its monthly spend limit of $50.00."
  }
}

Limits reset automatically at the start of each month. This is the single most effective guard against a runaway loop or a leaked key — set one on every key that leaves your machine.

Usage reporting

The Usage page in the Dashboard provides:

  • Requests, tokens and cost grouped by date, key and model.
  • Per-request detail: timestamp, model, input / cached / output tokens, and the cost of that request.
  • CSV export for internal chargeback.

Cost-saving techniques

TechniqueEffect
Match model to task difficultyClassification, extraction and reformatting on a small model often costs under a tenth of a flagship.
Keep prompt prefixes stablePut system messages and few-shot examples first and leave them unchanged to hit cached input.
Set max_tokensOutput is the priciest category; a cap prevents runaway responses.
Trim conversation historySummarise long threads instead of accumulating messages indefinitely.
Keep the tool list shortEvery tool definition consumes input tokens.
Downscale imagesImages convert to input tokens by resolution; oversized images are pure waste.

Common questions

  • Are failed requests billed? Requests rejected before reaching the upstream — 401, 402, 403, 404, 429 — are not billed. If generation had already begun before failure (upstream_error), the tokens produced are billed.
  • Is an aborted stream billed? Yes. Tokens the upstream already generated are billed.
  • Do credits expire? No.