Skip to content

Chat Completions

The full request parameters, message format and response shape for POST /v1/chat/completions.

Chat Completions

POST https://api.alphacurve.io/v1/chat/completions

This is the core endpoint. Send a list of messages, get a completion back. The request and response shapes are compatible with the OpenAI Chat Completions API.

Request parameters

FieldTypeDescription
model *stringModel ID, e.g. openai/gpt-4o, anthropic/claude-sonnet-4-5, google/gemini-2.5-pro. See Models.
messages *arrayConversation so far. Each item has a role (system | user | assistant | tool) and content.
streambooleanStream the completion as server-sent events. Defaults to false.
stream_optionsobjectStreaming options, e.g. {"include_usage": true} to receive usage in the final chunk.
temperaturenumberSampling temperature, 0–2. Lower is more deterministic. Defaults to the upstream model's default.
top_pnumberNucleus sampling. Use this or temperature, not both.
max_tokensintegerMaximum tokens to generate.
stopstring | arrayUp to 4 sequences that stop generation.
nintegerNumber of completions to generate. Defaults to 1.
presence_penaltynumber−2.0 to 2.0. Encourages new topics.
frequency_penaltynumber−2.0 to 2.0. Discourages repetition.
seedintegerBest-effort deterministic sampling. Not guaranteed.
toolsarrayTool definitions the model may call. See Tools.
tool_choicestring | objectauto, none, required, or a specific tool.
response_formatobject{"type": "json_object"} forces valid JSON output.
userstringYour own end-user identifier. It is recorded in usage reporting.

Fields marked * are required. Parameters an upstream model does not support are safely ignored rather than rejected, so the same code runs across models.

Message format

content may be a string, or an array of content parts (used for vision):

{
  "messages": [
    { "role": "system", "content": "You are concise." },
    { "role": "user", "content": "What is the tallest mountain in Taiwan?" },
    { "role": "assistant", "content": "Yushan, at 3,952 metres." },
    { "role": "user", "content": "And the second tallest?" }
  ]
}
RolePurpose
systemSystem instructions. Best placed first in the array.
userEnd-user input.
assistantA previous model reply. When it carries tool_calls, the model is asking you to run a tool.
toolThe result of a tool execution. Must include tool_call_id.

Examples

curl

curl https://api.alphacurve.io/v1/chat/completions \
  -H "Authorization: Bearer $INFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-sonnet-4-5",
    "messages": [
      { "role": "system", "content": "You are concise." },
      { "role": "user", "content": "Explain routing in one sentence." }
    ],
    "temperature": 0.7,
    "max_tokens": 256
  }'

Python

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.alphacurve.io/v1",
    api_key=os.environ["INFERENCE_API_KEY"],
)

resp = client.chat.completions.create(
    model="anthropic/claude-sonnet-4-5",
    messages=[
        {"role": "system", "content": "You are concise."},
        {"role": "user", "content": "Explain routing in one sentence."},
    ],
    temperature=0.7,
    max_tokens=256,
)

print(resp.choices[0].message.content)
print(resp.usage.prompt_tokens, resp.usage.completion_tokens)

Node / TypeScript

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.alphacurve.io/v1",
  apiKey: process.env.INFERENCE_API_KEY!,
});

const resp = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-4-5",
  messages: [
    { role: "system", content: "You are concise." },
    { role: "user", content: "Explain routing in one sentence." },
  ],
  temperature: 0.7,
  max_tokens: 256,
});

console.log(resp.choices[0].message.content);

Response format

{
  "id": "chatcmpl-3f9a1c2b7d4e",
  "object": "chat.completion",
  "created": 1753776000,
  "model": "anthropic/claude-sonnet-4-5",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Routing sends each request to the most suitable model that is currently available."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 28,
    "completion_tokens": 21,
    "total_tokens": 49,
    "prompt_tokens_details": { "cached_tokens": 0 }
  }
}
FieldDescription
idIdentifier for this completion. Include it when reporting a problem.
modelThe model that actually served the request.
choices[].finish_reasonstop (natural end), length (hit max_tokens), tool_calls (model wants a tool run), content_filter (filtered upstream).
usage.prompt_tokensInput tokens.
usage.completion_tokensOutput tokens.
usage.prompt_tokens_details.cached_tokensInput tokens served from cache, billed at the lower cached rate.

usage is exactly what we bill on. See Pricing.

JSON output

resp = client.chat.completions.create(
    model="openai/gpt-4o",
    messages=[
        {"role": "system", "content": "Output JSON only."},
        {"role": "user", "content": "Convert 'Yushan 3952 m' into {name, height_m}."},
    ],
    response_format={"type": "json_object"},
)

import json
print(json.loads(resp.choices[0].message.content))