Skip to content

Rate limits

Per-key RPM/TPM and concurrency limits, 429 headers, and backoff and queueing strategy.

Rate limits

Rate limits are enforced per API key to protect shared capacity and contain runaway traffic.

Default limits

LimitDefaultNotes
Requests per minute (RPM)600Sliding window.
Tokens per minute (TPM)2,000,000Input plus output tokens combined.
Concurrent requests60Requests in flight; a stream counts from start to finish.

Actual limits depend on your account tier and are shown on the API Keys page in the Dashboard. Contact us with your traffic profile if you need more headroom.

GET /v1/models has its own, much looser limit, so it is safe to use as a health check.

Response headers

Every response reports your current quota state:

X-RateLimit-Limit-Requests: 600
X-RateLimit-Remaining-Requests: 583
X-RateLimit-Reset-Requests: 42
X-RateLimit-Limit-Tokens: 2000000
X-RateLimit-Remaining-Tokens: 1904220
X-RateLimit-Reset-Tokens: 42

The *-Reset-* values are seconds until the window resets.

When you exceed a limit

HTTP/1.1 429 Too Many Requests
Retry-After: 12
Content-Type: application/json

{
  "error": {
    "code": "rate_limited",
    "message": "Rate limit exceeded for this API key. Retry in 12 seconds."
  }
}

Retry-After is the recommended wait in seconds. When present, prefer it over your own computed backoff.

Backoff and retry (Python)

import os
import random
import time

from openai import OpenAI, APIStatusError

client = OpenAI(
    base_url="https://api.alphacurve.io/v1",
    api_key=os.environ["INFERENCE_API_KEY"],
)

def complete_with_backoff(**kwargs):
    for attempt in range(6):
        try:
            return client.chat.completions.create(**kwargs)
        except APIStatusError as err:
            if err.status_code != 429:
                raise
            retry_after = err.response.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else min(2 ** attempt, 30)
            time.sleep(delay + random.uniform(0, 0.5))  # jitter avoids synchronised retries
    raise RuntimeError("rate limited: retries exhausted")

Bounding concurrency (Node / TypeScript)

Rather than absorbing 429s, cap how many requests are in flight:

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.alphacurve.io/v1",
  apiKey: process.env.INFERENCE_API_KEY!,
  maxRetries: 5, // the SDK retries 429 / 5xx for you
});

async function mapWithConcurrency<T, R>(
  items: T[],
  limit: number,
  fn: (item: T) => Promise<R>,
): Promise<R[]> {
  const results: R[] = new Array(items.length);
  let cursor = 0;

  const workers = Array.from({ length: limit }, async () => {
    while (cursor < items.length) {
      const index = cursor++;
      results[index] = await fn(items[index]);
    }
  });

  await Promise.all(workers);
  return results;
}

const answers = await mapWithConcurrency(prompts, 16, async (prompt) => {
  const resp = await client.chat.completions.create({
    model: "openai/gpt-4o-mini",
    messages: [{ role: "user", content: prompt }],
  });
  return resp.choices[0].message.content;
});

Practical guidance

  • Always add jitter. Fixed backoff makes every client retry in lockstep, which makes throttling worse.
  • Give batch jobs their own key. Keeping bulk work separate stops it from starving interactive traffic.
  • Watch X-RateLimit-Remaining-*. Slowing down as you approach the limit is smoother than hitting the wall and backing off.
  • Use streaming. It does not raise throughput, but it sharply reduces perceived latency.
  • Rate limits and spend limits are different things. 429 is a traffic problem; 402 is a balance or limit problem. Neither affects the other.