Models & naming
GET /v1/models, the provider/model naming scheme, and how to pick and restrict models.
Models & naming
Naming scheme
Model IDs always take the form provider/model:
openai/gpt-4o
openai/gpt-4o-mini
anthropic/claude-sonnet-4-5
anthropic/claude-haiku-4-5
google/gemini-2.5-pro
google/gemini-2.5-flash
| Segment | Description |
|---|---|
provider | The upstream provider: openai, anthropic, google, and other OpenAI-compatible providers. |
model | The provider's own model name, kept as close to the original as possible so their docs still apply. |
Prefixing avoids name collisions and makes it obvious from the model string alone where a request is going.
Listing available models
GET https://api.alphacurve.io/v1/models
The response reflects what this key can currently use. If the key has a model allowlist, only allowlisted models appear.
curl
curl https://api.alphacurve.io/v1/models \
-H "Authorization: Bearer $INFERENCE_API_KEY"
Python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
for m in client.models.list().data:
print(m.id)
Node / TypeScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
});
for (const m of (await client.models.list()).data) {
console.log(m.id);
}
Response
{
"object": "list",
"data": [
{
"id": "openai/gpt-4o",
"object": "model",
"created": 1753776000,
"owned_by": "openai",
"context_length": 128000,
"capabilities": {
"streaming": true,
"tools": true,
"vision": true
},
"pricing": {
"input_per_1m": "2.50",
"cached_input_per_1m": "1.25",
"output_per_1m": "10.00",
"currency": "USD"
}
}
]
}
pricing is the effective rate you are charged, after platform markup, per 1M tokens. See Pricing.
Choosing a model
| Situation | Suggestion |
|---|---|
| High-volume classification, extraction, summarisation | Small fast models such as openai/gpt-4o-mini, google/gemini-2.5-flash, anthropic/claude-haiku-4-5. |
| Complex reasoning and code | Flagship models such as openai/gpt-4o, anthropic/claude-sonnet-4-5, google/gemini-2.5-pro. |
| Very long documents | Pick a model with a large context_length, e.g. the Gemini family. |
| Image understanding | Models with capabilities.vision set to true. See Vision. |
| Function calling | Models with capabilities.tools set to true. See Tools. |
A pragmatic approach: build the pipeline on a cheap model, then upgrade only the steps where quality falls short.
Restricting models
Set a model allowlist per key in the Dashboard to prevent accidental calls to expensive models.
- Calling a model that does not exist on the platform →
404 model_not_found - Calling a model that exists but is not on this key's allowlist →
403 model_not_allowed
Fallbacks
Upstream providers occasionally time out or return 5xx. Keep a fallback model in your application layer:
MODELS = ["anthropic/claude-sonnet-4-5", "openai/gpt-4o"]
def complete(messages):
last_error = None
for model in MODELS:
try:
return client.chat.completions.create(model=model, messages=messages)
except Exception as exc: # timeout or upstream_error
last_error = exc
raise last_error