Skip to content
Claude Opus 4.5, GPT-5.2 and Gemini 3 Pro are live→

One endpoint,every model

Inference is an OpenAI-compatible gateway for large language models. Point your base_url at us and the same SDK, the same key, reaches OpenAI, Anthropic, Google Gemini and more — without touching a line of your code.

Prepaid credits · no subscription · no contract · revoke keys anytime

base_url to change
1
upstream providers
3+
failover availability
99.9%
monthly fee
0
~/your-app
curl https://api.alphacurve.io/v1/chat/completions \  -H "Authorization: Bearer $INFERENCE_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "model": "claude-opus-4-5",    "messages": [      { "role": "user", "content": "Hello, Inference!" }    ],    "stream": true  }'
Response200 OK · routed to Anthropic · 412ms to first token

Hello! I'm answering you through the Inference gateway.

Every major upstream, already wired — new models need zero code from you

◎OpenAI
✳Anthropic
✦Google Gemini
◈DeepSeek
𝕏xAI
∞Meta Llama
▨Mistral
⏦Groq
❖Qwen
☾Moonshot

Any OpenAI-compatible provider goes live as a catalog row, never a code change.

Why Inference

Turn multi-provider LLM plumbing into a one-time setup

You need more than a proxy. Key governance, spend caps, automatic failover and cost observability all live in the same layer.

One endpoint, every model

Fully compatible with the OpenAI Chat Completions wire format, including streaming, tool calls and vision. Your existing openai SDK, LangChain or Vercel AI SDK just points here — swap the base_url and you're done.

  • OpenAI-compatible
  • SSE streaming
  • Tools / vision

Multi-provider routing with automatic failover

Stack several credentials per upstream and we rotate them by priority. On a 429, a timeout or an upstream outage the gateway retries the next healthy key inside the same request — your users never see the failure.

  • Credential pool
  • 429 auto-retry
  • Health tracking

Fine-grained API key governance

Scope each key to an allow-list of models, give it a monthly spend cap, tag it by environment or team, and revoke it instantly. Contractors, staging and production stay properly isolated.

  • Model allow-list
  • Monthly spend cap
  • Instant revoke

Real-time usage and cost observability

Every request records token usage, cost, latency and which upstream served it. Slice by key, model or day to find the feature that is burning budget — instead of discovering one lump sum at month end.

  • Token breakdown
  • P95 latency
  • Cost per key

Prepaid credits, billed by the token

Top up your balance through Lemon Squeezy and we deduct against real token usage in real time. No monthly fee, no per-seat pricing, no minimum spend — and credits don't expire.

  • No subscription
  • Real-time metering
  • Credits never expire
Migration cost

Moving over from OpenAI is a one-line diff

No new SDK, no adapter layer, no rewriting your prompt pipeline. Request and response shapes stay exactly the same.

− Before
client.py
from openai import OpenAIclient = OpenAI(    base_url="https://api.openai.com/v1",    api_key=os.environ["OPENAI_API_KEY"],)# locked to a single vendor's modelsresp = client.chat.completions.create(    model="gpt-5.2",    messages=messages,)

Locked to one vendor — every model swap is a code change

+ After
client.py
from openai import OpenAIclient = OpenAI(    base_url="https://api.alphacurve.io/v1",    api_key=os.environ["INFERENCE_API_KEY"],)# any model in the catalog, same callresp = client.chat.completions.create(    model="claude-opus-4-5",    messages=messages,)

One SDK for every model, plus key governance and usage reporting

Streaming, function calling, multi-turn conversations and system prompts all keep working unchanged.

https://api.alphacurve.io/v1
Model catalog

Transparent pricing, metered per token

Every model shares one endpoint and one key. Switching is a change to the model field in your request — not a new vendor account.

  • claude-opus-4-5Anthropic
    VisionToolsNew
    Context
    200K
    Input / 1M
    $5.00
    Output / 1M
    $25.00
  • claude-sonnet-4-5Anthropic
    VisionTools
    Context
    200K
    Input / 1M
    $3.00
    Output / 1M
    $15.00
  • gpt-5.2OpenAI
    VisionToolsNew
    Context
    400K
    Input / 1M
    $1.25
    Output / 1M
    $10.00
  • gpt-5.2-miniOpenAI
    ToolsFast
    Context
    400K
    Input / 1M
    $0.25
    Output / 1M
    $2.00
  • gemini-3-proGoogle
    VisionTools
    Context
    1M
    Input / 1M
    $1.25
    Output / 1M
    $10.00
  • gemini-3-flashGoogle
    VisionFast
    Context
    1M
    Input / 1M
    $0.10
    Output / 1M
    $0.40
  • deepseek-v3.2DeepSeek
    ToolsFast
    Context
    128K
    Input / 1M
    $0.28
    Output / 1M
    $0.42
  • grok-4.1xAI
    VisionTools
    Context
    256K
    Input / 1M
    $3.00
    Output / 1M
    $15.00

Prices in USD, billed on actual token usage. Figures here are illustrative — the catalog API is authoritative.

See the full catalog
Observability

See exactly where the money goes

The dashboard aggregates token usage, cost and latency from every request in real time — sliced by key and by model, with alerts before a budget or balance runs out.

Usage overview

Last 14 daysLive
Spend this period

Spend this period

$48.16

+12.4%

Tokens used

Tokens used

34.2M

+8.1%

Requests

Requests

126,480

+3.6%

P95 latency

P95 latency

1.24s

-9.2%

Daily spend (USD)

Last 14 days

Cost by model

  • claude-opus-4-542%
  • gpt-5.227%
  • gemini-3-pro18%
  • deepseek-v3.213%

Usage by key

  • prod-web$26.40
  • prod-worker$14.02
  • stagingNear cap$5.91
  • agency-demo$1.83
Pricing

Prepaid credits. No subscription, no contract.

Top up once through Lemon Squeezy and we meter against your balance in real time. Unused credits stay yours, and you can stop any time — there is nothing to cancel.

Starter

$10

For a proof of concept, a side project or validating the integration.

≈ 8M tokens on gemini-3-flash

Top up $10
Most popular

Builder

$50+$2.50 bonus

You have real users and need failover plus proper usage reporting.

≈ 17.5M tokens across a typical model mix

Top up $50

Scale

$200+$14.00 bonus

Multiple environments and team keys that need spend caps and cost attribution.

≈ 71M tokens across a typical model mix

Top up $200

Need a larger balance or an invoice? Talk to us

How billing works

Charged on tokens, nothing else
After each request we deduct the model's input and output rates multiplied by real token counts, resolved to micro-dollars.
No monthly or per-seat fee
No subscription, no contract, no minimum spend. A month with no calls costs you nothing.
Credits don't expire
Your balance stays put, and the dashboard can warn you — or auto top up — before it runs low.
Spend caps per key
Give each key a monthly cap and requests are refused past it — one runaway script can't drain the balance.
FAQ

Anything else on your mind?

If it isn't answered here, the docs go deeper — or just email us.

  • Yes. We implement the OpenAI Chat Completions API, so request and response payloads are identical — including SSE streaming, function calling and multimodal input. Point base_url at https://api.alphacurve.io/v1, use an Inference key, and your SDK, prompts and post-processing stay exactly as they are.

Get started

Wire it up in five minutes, never rewrite for a model swap again

Sign up, create your first key, and run the whole flow on trial credits before you top up. No credit card, no sales call.

Swap this one line and you're live

base_url = "https://api.alphacurve.io/v1"
  • No credit card
  • OpenAI-compatible
  • Revoke keys anytime