Streaming
Set stream: true to receive tokens over SSE, and how to get usage from a stream.
Streaming
Set stream to true and the completion arrives incrementally as Server-Sent Events (SSE). This matters for chat interfaces — users see text as it is generated rather than waiting for the whole reply.
Wire format
Each event is a data: line carrying a chat.completion.chunk JSON object. The stream terminates with data: [DONE].
data: {"id":"chatcmpl-3f9a","object":"chat.completion.chunk","created":1753776000,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-3f9a","object":"chat.completion.chunk","created":1753776000,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{"content":"Routing"},"finish_reason":null}]}
data: {"id":"chatcmpl-3f9a","object":"chat.completion.chunk","created":1753776000,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{"content":" sends"},"finish_reason":null}]}
data: {"id":"chatcmpl-3f9a","object":"chat.completion.chunk","created":1753776000,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]
Two differences from a non-streamed response: content lives in choices[0].delta.content rather than choices[0].message.content, and the first chunk typically carries only role with empty content.
curl
Add --no-buffer (or -N), otherwise curl waits for the whole body before printing:
curl --no-buffer https://api.alphacurve.io/v1/chat/completions \
-H "Authorization: Bearer $INFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o-mini",
"messages": [{ "role": "user", "content": "Write a haiku about routing." }],
"stream": true
}'
Python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
stream = client.chat.completions.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Write a haiku about routing."}],
stream=True,
stream_options={"include_usage": True},
)
full = []
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
piece = chunk.choices[0].delta.content
full.append(piece)
print(piece, end="", flush=True)
if chunk.usage: # final chunk
print()
print("usage:", chunk.usage)
print("".join(full))
Node / TypeScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
});
const stream = await client.chat.completions.create({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "Write a haiku about routing." }],
stream: true,
stream_options: { include_usage: true },
});
let full = "";
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta?.content ?? "";
full += delta;
process.stdout.write(delta);
if (chunk.usage) console.log("\nusage:", chunk.usage);
}
Parsing SSE without an SDK
const res = await fetch("https://api.alphacurve.io/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.INFERENCE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "Hello" }],
stream: true,
}),
});
const reader = res.body!.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() ?? "";
for (const line of lines) {
if (!line.startsWith("data: ")) continue;
const payload = line.slice(6).trim();
if (payload === "[DONE]") continue;
const chunk = JSON.parse(payload);
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
}
Usage and billing while streaming
Pass stream_options: {"include_usage": true} and the final chunk carries usage (with an empty choices array).
If you disconnect mid-stream, tokens already generated are still billed — the upstream provider has already produced them. Use max_tokens to bound cost.
Notes
- Once streaming begins the HTTP status is already
200. If an upstream failure occurs mid-stream, the error is delivered as adata:event containing the standard error object (see Errors). - If Nginx sits in front of your service, set
proxy_buffering off;or the stream will be buffered. - On Vercel / Cloudflare, use an edge or streaming runtime and forward the
Responsebody directly. - Under streaming, tool calls arrive as
delta.tool_callsfragments. Accumulateargumentsstrings byindexbefore callingJSON.parse.