串流(Streaming)
設定 stream: true 以 SSE 逐 token 接收回覆,並取得串流中的用量統計。
串流(Streaming)
把 stream 設為 true,回覆就會以 Server-Sent Events(SSE) 逐段送出。這對聊天介面很重要——使用者不必等到整段生成完才看到內容。
傳輸格式
每一筆事件是一行 data: ,後面接一個 chat.completion.chunk JSON 物件;串流以 data: [DONE] 結束。
data: {"id":"chatcmpl-3f9a","object":"chat.completion.chunk","created":1753776000,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-3f9a","object":"chat.completion.chunk","created":1753776000,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{"content":"路由"},"finish_reason":null}]}
data: {"id":"chatcmpl-3f9a","object":"chat.completion.chunk","created":1753776000,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{"content":"就是"},"finish_reason":null}]}
data: {"id":"chatcmpl-3f9a","object":"chat.completion.chunk","created":1753776000,"model":"openai/gpt-4o-mini","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]
與非串流回應的差異:內容在 choices[0].delta.content 而不是 choices[0].message.content,而且第一個 chunk 通常只帶 role、內容為空字串。
curl
加上 --no-buffer(或 -N),否則 curl 會等到整段結束才輸出:
curl --no-buffer https://api.alphacurve.io/v1/chat/completions \
-H "Authorization: Bearer $INFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o-mini",
"messages": [{ "role": "user", "content": "寫一首關於路由的俳句。" }],
"stream": true
}'
Python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
stream = client.chat.completions.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "寫一首關於路由的俳句。"}],
stream=True,
stream_options={"include_usage": True},
)
full = []
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
piece = chunk.choices[0].delta.content
full.append(piece)
print(piece, end="", flush=True)
if chunk.usage: # 最後一個 chunk
print()
print("usage:", chunk.usage)
print("".join(full))
Node / TypeScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
});
const stream = await client.chat.completions.create({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "寫一首關於路由的俳句。" }],
stream: true,
stream_options: { include_usage: true },
});
let full = "";
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta?.content ?? "";
full += delta;
process.stdout.write(delta);
if (chunk.usage) console.log("\nusage:", chunk.usage);
}
不用 SDK,直接處理 SSE
const res = await fetch("https://api.alphacurve.io/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.INFERENCE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "Hello" }],
stream: true,
}),
});
const reader = res.body!.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop() ?? "";
for (const line of lines) {
if (!line.startsWith("data: ")) continue;
const payload = line.slice(6).trim();
if (payload === "[DONE]") continue;
const chunk = JSON.parse(payload);
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
}
串流中的用量與計費
加上 stream_options: {"include_usage": true},最後一個 chunk 會帶 usage(此時 choices 為空陣列)。
即使你在中途中斷連線,已產生的 token 仍會計費——上游已經算過了。想控制成本請用 max_tokens。
注意事項
- 串流開始後 HTTP 狀態碼已是
200。若上游在串流途中失敗,錯誤會以一筆data:事件送出,內容為標準錯誤物件(見 錯誤)。 - 若你的服務前面有 Nginx,記得設定
proxy_buffering off;,否則串流會被緩衝。 - 在 Vercel / Cloudflare 等平台請使用 Edge / streaming runtime,並直接把
Responsebody 轉發出去。 - Tool calling 在串流下會以
delta.tool_calls分段送出arguments,需要自行按index累積字串後再JSON.parse。