Skip to content

Models & naming

GET /v1/models, the provider/model naming scheme, and how to pick and restrict models.

Models & naming

Naming scheme

Model IDs always take the form provider/model:

openai/gpt-4o
openai/gpt-4o-mini
anthropic/claude-sonnet-4-5
anthropic/claude-haiku-4-5
google/gemini-2.5-pro
google/gemini-2.5-flash
SegmentDescription
providerThe upstream provider: openai, anthropic, google, and other OpenAI-compatible providers.
modelThe provider's own model name, kept as close to the original as possible so their docs still apply.

Prefixing avoids name collisions and makes it obvious from the model string alone where a request is going.

Listing available models

GET https://api.alphacurve.io/v1/models

The response reflects what this key can currently use. If the key has a model allowlist, only allowlisted models appear.

curl

curl https://api.alphacurve.io/v1/models \
  -H "Authorization: Bearer $INFERENCE_API_KEY"

Python

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.alphacurve.io/v1",
    api_key=os.environ["INFERENCE_API_KEY"],
)

for m in client.models.list().data:
    print(m.id)

Node / TypeScript

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.alphacurve.io/v1",
  apiKey: process.env.INFERENCE_API_KEY!,
});

for (const m of (await client.models.list()).data) {
  console.log(m.id);
}

Response

{
  "object": "list",
  "data": [
    {
      "id": "openai/gpt-4o",
      "object": "model",
      "created": 1753776000,
      "owned_by": "openai",
      "context_length": 128000,
      "capabilities": {
        "streaming": true,
        "tools": true,
        "vision": true
      },
      "pricing": {
        "input_per_1m": "2.50",
        "cached_input_per_1m": "1.25",
        "output_per_1m": "10.00",
        "currency": "USD"
      }
    }
  ]
}

pricing is the effective rate you are charged, after platform markup, per 1M tokens. See Pricing.

Choosing a model

SituationSuggestion
High-volume classification, extraction, summarisationSmall fast models such as openai/gpt-4o-mini, google/gemini-2.5-flash, anthropic/claude-haiku-4-5.
Complex reasoning and codeFlagship models such as openai/gpt-4o, anthropic/claude-sonnet-4-5, google/gemini-2.5-pro.
Very long documentsPick a model with a large context_length, e.g. the Gemini family.
Image understandingModels with capabilities.vision set to true. See Vision.
Function callingModels with capabilities.tools set to true. See Tools.

A pragmatic approach: build the pipeline on a cheap model, then upgrade only the steps where quality falls short.

Restricting models

Set a model allowlist per key in the Dashboard to prevent accidental calls to expensive models.

  • Calling a model that does not exist on the platform → 404 model_not_found
  • Calling a model that exists but is not on this key's allowlist → 403 model_not_allowed

Fallbacks

Upstream providers occasionally time out or return 5xx. Keep a fallback model in your application layer:

MODELS = ["anthropic/claude-sonnet-4-5", "openai/gpt-4o"]

def complete(messages):
    last_error = None
    for model in MODELS:
        try:
            return client.chat.completions.create(model=model, messages=messages)
        except Exception as exc:   # timeout or upstream_error
            last_error = exc
    raise last_error