Introduction
Inference is an OpenAI-compatible LLM API gateway that gives you one key for OpenAI, Anthropic and Google Gemini models.
Introduction
Inference is an OpenAI-compatible LLM API gateway. With a single API key and a single base URL you can reach models from multiple upstream providers: OpenAI-compatible providers, Anthropic, and Google Gemini.
If you have ever written against the OpenAI API, you already know how to use Inference — point base_url at our domain, swap in an sk-inf-... key, and the rest of your code stays exactly the same.
Why Inference
| Benefit | Description |
|---|---|
| One integration | A single OpenAI-compatible request/response shape across every upstream provider. |
| One bill | Prepaid credit balance, charged by token usage. No separate accounts with each provider. |
| Cost control | Every key can carry its own monthly spend limit and model allowlist. |
| Observability | The Dashboard shows model, token usage and cost for every request. |
| Standard features | Streaming (SSE), tools / function calling, and vision (image input). |
Endpoints
The base URL is always:
https://api.alphacurve.io/v1
| Method | Path | Purpose |
|---|---|---|
POST | /v1/chat/completions | Chat completions, with streaming, tools and vision. |
GET | /v1/models | List the models currently available to your key. |
WSS | /ws/…BidiGenerateContent | Gemini Live realtime audio, bidirectional WebSocket. See below. |
Realtime audio (Gemini Live)
Live models are not OpenAI-shaped and cannot be called through /v1/chat/completions. They are a stateful, bidirectional WebSocket conversation, speaking exactly the protocol Google's Live API speaks:
wss://api.alphacurve.io/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent?key=sk-inf-...
Existing Google Live client code works here by changing two things — the host and the key. Nothing else. The key goes in the query string because that is what the protocol does (a browser cannot set headers on a WebSocket handshake); non-browser clients should send Authorization: Bearer instead, keeping the secret out of URLs and out of every log along the way.
Audio comes back as 24 kHz 16-bit mono PCM in serverContent.modelTurn.parts[].inlineData.data (base64). It has no container, so wrap it in a WAV header before playing.
Redaction does not apply on this path. Neither
X-Inference-Redactnor a model's redaction setting has any effect on a realtime session — there is no text in a PCM stream to scan and replace, so audio reaches the upstream as sent. If redaction is why you chose this platform, do not send sensitive content over realtime.
30-second tour
curl https://api.alphacurve.io/v1/chat/completions \
-H "Authorization: Bearer $INFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o-mini",
"messages": [
{ "role": "user", "content": "Explain what an API gateway is in one sentence." }
]
}'
Next steps
- Quickstart — send your first request in three steps.
- Authentication — key format, management and security guidance.
- Chat Completions — the full parameter and response reference.
- Migration guide — moving over from OpenAI or Anthropic.