Chat Completions
The full request parameters, message format and response shape for POST /v1/chat/completions.
Chat Completions
POST https://api.alphacurve.io/v1/chat/completions
This is the core endpoint. Send a list of messages, get a completion back. The request and response shapes are compatible with the OpenAI Chat Completions API.
Request parameters
| Field | Type | Description |
|---|---|---|
model * | string | Model ID, e.g. openai/gpt-4o, anthropic/claude-sonnet-4-5, google/gemini-2.5-pro. See Models. |
messages * | array | Conversation so far. Each item has a role (system | user | assistant | tool) and content. |
stream | boolean | Stream the completion as server-sent events. Defaults to false. |
stream_options | object | Streaming options, e.g. {"include_usage": true} to receive usage in the final chunk. |
temperature | number | Sampling temperature, 0–2. Lower is more deterministic. Defaults to the upstream model's default. |
top_p | number | Nucleus sampling. Use this or temperature, not both. |
max_tokens | integer | Maximum tokens to generate. |
stop | string | array | Up to 4 sequences that stop generation. |
n | integer | Number of completions to generate. Defaults to 1. |
presence_penalty | number | −2.0 to 2.0. Encourages new topics. |
frequency_penalty | number | −2.0 to 2.0. Discourages repetition. |
seed | integer | Best-effort deterministic sampling. Not guaranteed. |
tools | array | Tool definitions the model may call. See Tools. |
tool_choice | string | object | auto, none, required, or a specific tool. |
response_format | object | {"type": "json_object"} forces valid JSON output. |
user | string | Your own end-user identifier. It is recorded in usage reporting. |
Fields marked * are required. Parameters an upstream model does not support are safely ignored rather than rejected, so the same code runs across models.
Message format
content may be a string, or an array of content parts (used for vision):
{
"messages": [
{ "role": "system", "content": "You are concise." },
{ "role": "user", "content": "What is the tallest mountain in Taiwan?" },
{ "role": "assistant", "content": "Yushan, at 3,952 metres." },
{ "role": "user", "content": "And the second tallest?" }
]
}
| Role | Purpose |
|---|---|
system | System instructions. Best placed first in the array. |
user | End-user input. |
assistant | A previous model reply. When it carries tool_calls, the model is asking you to run a tool. |
tool | The result of a tool execution. Must include tool_call_id. |
Examples
curl
curl https://api.alphacurve.io/v1/chat/completions \
-H "Authorization: Bearer $INFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4-5",
"messages": [
{ "role": "system", "content": "You are concise." },
{ "role": "user", "content": "Explain routing in one sentence." }
],
"temperature": 0.7,
"max_tokens": 256
}'
Python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
resp = client.chat.completions.create(
model="anthropic/claude-sonnet-4-5",
messages=[
{"role": "system", "content": "You are concise."},
{"role": "user", "content": "Explain routing in one sentence."},
],
temperature=0.7,
max_tokens=256,
)
print(resp.choices[0].message.content)
print(resp.usage.prompt_tokens, resp.usage.completion_tokens)
Node / TypeScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
});
const resp = await client.chat.completions.create({
model: "anthropic/claude-sonnet-4-5",
messages: [
{ role: "system", content: "You are concise." },
{ role: "user", content: "Explain routing in one sentence." },
],
temperature: 0.7,
max_tokens: 256,
});
console.log(resp.choices[0].message.content);
Response format
{
"id": "chatcmpl-3f9a1c2b7d4e",
"object": "chat.completion",
"created": 1753776000,
"model": "anthropic/claude-sonnet-4-5",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Routing sends each request to the most suitable model that is currently available."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 28,
"completion_tokens": 21,
"total_tokens": 49,
"prompt_tokens_details": { "cached_tokens": 0 }
}
}
| Field | Description |
|---|---|
id | Identifier for this completion. Include it when reporting a problem. |
model | The model that actually served the request. |
choices[].finish_reason | stop (natural end), length (hit max_tokens), tool_calls (model wants a tool run), content_filter (filtered upstream). |
usage.prompt_tokens | Input tokens. |
usage.completion_tokens | Output tokens. |
usage.prompt_tokens_details.cached_tokens | Input tokens served from cache, billed at the lower cached rate. |
usage is exactly what we bill on. See Pricing.
JSON output
resp = client.chat.completions.create(
model="openai/gpt-4o",
messages=[
{"role": "system", "content": "Output JSON only."},
{"role": "user", "content": "Convert 'Yushan 3952 m' into {name, height_m}."},
],
response_format={"type": "json_object"},
)
import json
print(json.loads(resp.choices[0].message.content))