SDKs & language integrations
Connect via the official OpenAI SDKs, raw HTTP, or common frameworks such as LangChain and the Vercel AI SDK.
SDKs & language integrations
Inference is an OpenAI-compatible API, so any OpenAI client that lets you override the base URL works out of the box. There is no proprietary SDK to install.
The common configuration:
base URL : https://api.alphacurve.io/v1
API key : sk-inf-...
curl
export INFERENCE_API_KEY="sk-inf-xxxxxxxxxxxxxxxxxxxxxxxx"
curl https://api.alphacurve.io/v1/chat/completions \
-H "Authorization: Bearer $INFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o-mini",
"messages": [{ "role": "user", "content": "Hello" }]
}'
Python — openai SDK
pip install openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
timeout=60.0,
max_retries=3,
)
resp = client.chat.completions.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)
Async variant:
import asyncio
import os
from openai import AsyncOpenAI
client = AsyncOpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
async def main():
resp = await client.chat.completions.create(
model="google/gemini-2.5-flash",
messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)
asyncio.run(main())
Python — without an SDK
import os
import requests
resp = requests.post(
"https://api.alphacurve.io/v1/chat/completions",
headers={
"Authorization": f"Bearer {os.environ['INFERENCE_API_KEY']}",
"Content-Type": "application/json",
},
json={
"model": "openai/gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello"}],
},
timeout=60,
)
resp.raise_for_status()
print(resp.json()["choices"][0]["message"]["content"])
Node / TypeScript — openai SDK
npm install openai
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
timeout: 60_000,
maxRetries: 3,
});
const resp = await client.chat.completions.create({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "Hello" }],
});
console.log(resp.choices[0].message.content);
Node / TypeScript — plain fetch
const res = await fetch("https://api.alphacurve.io/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.INFERENCE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "openai/gpt-4o-mini",
messages: [{ role: "user", content: "Hello" }],
}),
});
if (!res.ok) {
const { error } = await res.json();
throw new Error(`${error.code}: ${error.message}`);
}
const data = await res.json();
console.log(data.choices[0].message.content);
Vercel AI SDK
npm install ai @ai-sdk/openai-compatible
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { streamText } from "ai";
const inference = createOpenAICompatible({
name: "inference",
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
});
const result = streamText({
model: inference("anthropic/claude-sonnet-4-5"),
prompt: "Explain routing in one sentence.",
});
for await (const chunk of result.textStream) {
process.stdout.write(chunk);
}
LangChain (Python)
pip install langchain-openai
import os
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="anthropic/claude-sonnet-4-5",
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
print(llm.invoke("Explain routing in one sentence.").content)
LlamaIndex (Python)
import os
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
model="google/gemini-2.5-pro",
api_base="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
is_chat_model=True,
)
print(llm.complete("Explain routing in one sentence."))
Go
Go has no single standard OpenAI SDK; the standard library is enough:
package main
import (
"bytes"
"encoding/json"
"fmt"
"net/http"
"os"
)
func main() {
body, _ := json.Marshal(map[string]any{
"model": "openai/gpt-4o-mini",
"messages": []map[string]string{
{"role": "user", "content": "Hello"},
},
})
req, _ := http.NewRequest("POST",
"https://api.alphacurve.io/v1/chat/completions",
bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFERENCE_API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
var out struct {
Choices []struct {
Message struct {
Content string `json:"content"`
} `json:"message"`
} `json:"choices"`
}
json.NewDecoder(resp.Body).Decode(&out)
fmt.Println(out.Choices[0].Message.Content)
}
Recommended client settings
| Setting | Suggested value | Why |
|---|---|---|
timeout | 60–120 seconds | Long outputs and reasoning models take time; default timeouts are often too short. |
max_retries | 3–5 | Built-in SDK retries cover 429 and 5xx. |
| base URL | Read from an environment variable | Lets you switch environments, or point back at a provider directly, without code changes. |