計費方式
預付 credits、每 1M tokens 計價、input/output/cached 分開計費,以及餘額與花費上限。
計費方式
Inference 採預付儲值制。你先在 Dashboard 儲值 credits,每次請求結束後依實際 token 用量從餘額扣款。沒有月租費,也沒有最低消費。
計價單位
所有價格都以每 1,000,000 tokens(1M tokens)美元表示,並分成三種:
| 類別 | 說明 |
|---|---|
| Input | 你送出的 prompt token,包含 system 訊息、對話歷史、工具定義與圖片換算出的 token。 |
| Cached input | 命中上游 prompt cache 的輸入 token,單價明顯低於一般 input。 |
| Output | 模型生成的 token,通常是三者中最貴的。 |
單次請求成本的計算方式:
cost = (input_tokens / 1_000_000) × input_per_1m
+ (cached_input_tokens / 1_000_000) × cached_input_per_1m
+ (output_tokens / 1_000_000) × output_per_1m
平台加成(markup)
我們在上游供應商的成本上加一個固定比例的平台加成,用來支撐 gateway 的維運、備援與帳務。GET /v1/models 回傳的 pricing 已經包含加成,也就是你實際被扣款的單價——不會再有其他隱藏費用。
各模型的加成比例一致且透明,隨時可在 Dashboard 的 Pricing 頁面查看。
價格範例
以下為說明用的示意數字,實際單價一律以 GET /v1/models 與 Dashboard 為準。
| 模型 | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
openai/gpt-4o | $2.50 | $1.25 | $10.00 |
openai/gpt-4o-mini | $0.15 | $0.075 | $0.60 |
anthropic/claude-sonnet-4-5 | $3.00 | $0.30 | $15.00 |
google/gemini-2.5-pro | $1.25 | $0.31 | $10.00 |
google/gemini-2.5-flash | $0.30 | $0.075 | $2.50 |
從回應算出成本
每一次回應的 usage 就是計費依據:
{
"usage": {
"prompt_tokens": 12000,
"completion_tokens": 800,
"total_tokens": 12800,
"prompt_tokens_details": { "cached_tokens": 9000 }
}
}
注意 prompt_tokens 已包含 cached_tokens,計算時要先扣掉:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
PRICING = { # 每 1M tokens 美元;正式使用請改為呼叫 /v1/models 取得
"openai/gpt-4o": {"input": 2.50, "cached_input": 1.25, "output": 10.00},
}
resp = client.chat.completions.create(
model="openai/gpt-4o",
messages=[{"role": "user", "content": "Hello"}],
)
u = resp.usage
cached = (u.prompt_tokens_details.cached_tokens if u.prompt_tokens_details else 0) or 0
fresh_input = u.prompt_tokens - cached
p = PRICING[resp.model]
cost = (
fresh_input / 1_000_000 * p["input"]
+ cached / 1_000_000 * p["cached_input"]
+ u.completion_tokens / 1_000_000 * p["output"]
)
print(f"cost: ${cost:.6f}")
同樣的計算在 Node:
const u = resp.usage!;
const cached = u.prompt_tokens_details?.cached_tokens ?? 0;
const freshInput = u.prompt_tokens - cached;
const cost =
(freshInput / 1_000_000) * 2.5 +
(cached / 1_000_000) * 1.25 +
(u.completion_tokens / 1_000_000) * 10.0;
console.log(`cost: $${cost.toFixed(6)}`);
餘額不足
發送請求時若餘額不足以支付,會直接被拒絕:
HTTP/1.1 402 Payment Required
Content-Type: application/json
{
"error": {
"code": "insufficient_credit",
"message": "Your credit balance is insufficient for this request."
}
}
到 Dashboard 的 Billing 儲值後即可恢復。建議在 Dashboard 設定低餘額通知,避免生產環境突然中斷。
每月花費上限
每把金鑰都可以獨立設定每月花費上限(美元)。當月累計成本達到上限後:
HTTP/1.1 402 Payment Required
Content-Type: application/json
{
"error": {
"code": "spend_limit_reached",
"message": "This API key has reached its monthly spend limit of $50.00."
}
}
上限在每月初自動重置。這是防止程式失控或金鑰外洩造成損失的最有效手段——請務必為每把對外的金鑰設定上限。
用量查詢
Dashboard 的 Usage 頁面提供:
- 依日期、金鑰、模型分組的請求數、token 數與成本。
- 逐筆請求明細:時間、模型、input / cached / output token 數、該筆成本。
- 可匯出 CSV,方便做內部分帳。
省錢技巧
| 做法 | 效果 |
|---|---|
| 依任務難度選模型 | 分類、抽取、格式轉換用小模型,成本常常只有旗艦模型的 1/10 以下。 |
| 讓 prompt 前綴保持穩定 | system 訊息與少樣本範例放在最前面且不變動,較容易命中 cached input。 |
設定 max_tokens | output 最貴,設定上限可避免異常長的回覆。 |
| 精簡對話歷史 | 長對話請摘要壓縮,而不是無限累積訊息。 |
| 控制工具數量 | 每個工具定義都佔用 input token。 |
| 圖片先縮放 | 圖片按解析度換算成 input token,過大的圖只是浪費。 |
常見問題
- 失敗的請求會計費嗎? 401、402、403、404、429 這類在送到上游之前就被擋下的請求不計費。若上游已經開始生成才失敗(
upstream_error),已產生的 token 會被計費。 - 串流中途中斷會計費嗎? 會。上游已經產生的 token 仍然計費。
- credits 會過期嗎? 不會。