Rate limits
Per-key RPM/TPM and concurrency limits, 429 headers, and backoff and queueing strategy.
Rate limits
Rate limits are enforced per API key to protect shared capacity and contain runaway traffic.
Default limits
| Limit | Default | Notes |
|---|---|---|
| Requests per minute (RPM) | 600 | Sliding window. |
| Tokens per minute (TPM) | 2,000,000 | Input plus output tokens combined. |
| Concurrent requests | 60 | Requests in flight; a stream counts from start to finish. |
Actual limits depend on your account tier and are shown on the API Keys page in the Dashboard. Contact us with your traffic profile if you need more headroom.
GET /v1/models has its own, much looser limit, so it is safe to use as a health check.
Response headers
Every response reports your current quota state:
X-RateLimit-Limit-Requests: 600
X-RateLimit-Remaining-Requests: 583
X-RateLimit-Reset-Requests: 42
X-RateLimit-Limit-Tokens: 2000000
X-RateLimit-Remaining-Tokens: 1904220
X-RateLimit-Reset-Tokens: 42
The *-Reset-* values are seconds until the window resets.
When you exceed a limit
HTTP/1.1 429 Too Many Requests
Retry-After: 12
Content-Type: application/json
{
"error": {
"code": "rate_limited",
"message": "Rate limit exceeded for this API key. Retry in 12 seconds."
}
}
Retry-After is the recommended wait in seconds. When present, prefer it over your own computed backoff.
Backoff and retry (Python)
import os
import random
import time
from openai import OpenAI, APIStatusError
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
def complete_with_backoff(**kwargs):
for attempt in range(6):
try:
return client.chat.completions.create(**kwargs)
except APIStatusError as err:
if err.status_code != 429:
raise
retry_after = err.response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else min(2 ** attempt, 30)
time.sleep(delay + random.uniform(0, 0.5)) # jitter avoids synchronised retries
raise RuntimeError("rate limited: retries exhausted")
Bounding concurrency (Node / TypeScript)
Rather than absorbing 429s, cap how many requests are in flight:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
maxRetries: 5, // the SDK retries 429 / 5xx for you
});
async function mapWithConcurrency<T, R>(
items: T[],
limit: number,
fn: (item: T) => Promise<R>,
): Promise<R[]> {
const results: R[] = new Array(items.length);
let cursor = 0;
const workers = Array.from({ length: limit }, async () => {
while (cursor < items.length) {
const index = cursor++;
results[index] = await fn(items[index]);
}
});
await Promise.all(workers);
return results;
}
const answers = await mapWithConcurrency(prompts, 16, async (prompt) => {
const resp = await client.chat.completions.create({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: prompt }],
});
return resp.choices[0].message.content;
});
Practical guidance
- Always add jitter. Fixed backoff makes every client retry in lockstep, which makes throttling worse.
- Give batch jobs their own key. Keeping bulk work separate stops it from starving interactive traffic.
- Watch
X-RateLimit-Remaining-*. Slowing down as you approach the limit is smoother than hitting the wall and backing off. - Use streaming. It does not raise throughput, but it sharply reduces perceived latency.
- Rate limits and spend limits are different things. 429 is a traffic problem; 402 is a balance or limit problem. Neither affects the other.