FAQ
Common questions about compatibility, models, billing, data handling and troubleshooting.
FAQ
General
How is Inference different from calling OpenAI or Anthropic directly? The difference is integration cost, not the models themselves. One key, one base URL and one codebase reach several providers; billing is consolidated; usage and cost live in a single Dashboard; and switching models is a one-string change.
Do I need a proprietary SDK? No. Any OpenAI client that supports a custom base URL works. See SDKs & language integrations.
Do you support every OpenAI endpoint?
Today we serve POST /v1/chat/completions and GET /v1/models. Unsupported parameters are safely ignored rather than rejected.
Models
Which models are available?
Call GET /v1/models. The response reflects what your key can currently use.
Why does gpt-4o return 404?
Model IDs need the provider prefix: openai/gpt-4o. See Models & naming.
Can I restrict a key to specific models?
Yes. Set a model allowlist on the key in the Dashboard. Anything outside it returns 403 model_not_allowed.
Will you silently substitute a different model? No. We route to the model you name. Implement fallbacks in your application layer — see Models.
Billing
How is pricing calculated? Per 1M tokens, with input, cached input and output priced separately, platform markup included. See Pricing & billing.
What happens when my balance runs out?
The API returns 402 insufficient_credit. Top up in the Dashboard to resume, and turn on low-balance notifications beforehand.
How do I stop overspending?
Set a monthly spend limit on each key. Once reached, requests return 402 spend_limit_reached, and the limit resets the following month.
Are failed requests billed? Requests rejected before reaching the upstream (401, 402, 403, 404, 429) are not billed. If generation had already begun, the tokens produced are billed.
Do credits expire? No.
Features
Is streaming supported?
Yes. Set stream: true and consume standard SSE. See Streaming.
Is function calling supported? Yes, in the unified OpenAI format across providers. See Tools.
Is image input supported? Yes, on vision-capable models. See Vision.
What about embeddings, image generation or audio? Not currently. We are focused on chat completions.
Data & privacy
Do you store my prompts? We retain the metadata required for billing and audit — timestamp, model, token counts and cost. Request and response content is not used to train models. See the privacy policy for the full terms.
Where do my requests go?
To the upstream provider named in your model field. That provider's policies also apply to your data.
Troubleshooting
I keep getting 401.
Check the header is Authorization: Bearer sk-inf-..., that the key is complete and free of stray whitespace, and that it is not disabled. curl -i https://api.alphacurve.io/v1/models is the quickest check.
I'm getting 429.
You have exceeded the key's rate limit. Back off per Retry-After and lower concurrency. See Rate limits.
I'm getting 502 upstream_error.
The upstream provider failed or timed out. Safe to retry, or fail over to another model.
Streaming arrives all at once instead of token by token.
Usually a proxy buffering the response. Add --no-buffer to curl, set proxy_buffering off; in Nginx, and use a streaming or edge runtime on serverless platforms.
My response is cut off.
Check finish_reason. length means you hit max_tokens — raise it.
How do I report a problem?
Send us the response id, the error code, the timestamp and the model used.