Vision (image input)
Pass image URLs or base64 data through image_url content parts so models can read images.
Vision (image input)
Vision-capable models can read images. Change content from a string to an array of content parts, mixing text and image_url entries.
Content part format
{
"role": "user",
"content": [
{ "type": "text", "text": "What is in this image?" },
{
"type": "image_url",
"image_url": {
"url": "https://example.com/photo.jpg",
"detail": "auto"
}
}
]
}
| Field | Description |
|---|---|
image_url.url | A publicly reachable HTTPS URL, or a data: base64 URI. |
image_url.detail | auto (default), low, high. low uses fewer tokens and costs less; high preserves detail. Some upstream models ignore this field. |
Supported formats: PNG, JPEG, WebP, and non-animated GIF.
Passing a URL
curl
curl https://api.alphacurve.io/v1/chat/completions \
-H "Authorization: Bearer $INFERENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o",
"messages": [{
"role": "user",
"content": [
{ "type": "text", "text": "What is in this image?" },
{ "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } }
]
}]
}'
Python
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
resp = client.chat.completions.create(
model="openai/gpt-4o",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url",
"image_url": {"url": "https://example.com/photo.jpg"}},
],
}],
)
print(resp.choices[0].message.content)
Node / TypeScript
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
});
const resp = await client.chat.completions.create({
model: "google/gemini-2.5-pro",
messages: [{
role: "user",
content: [
{ type: "text", text: "What is in this image?" },
{ type: "image_url", image_url: { url: "https://example.com/photo.jpg" } },
],
}],
});
console.log(resp.choices[0].message.content);
Passing a local file (base64)
When the image is not on the public internet, use a data: URI.
Python
import base64
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.alphacurve.io/v1",
api_key=os.environ["INFERENCE_API_KEY"],
)
with open("receipt.png", "rb") as f:
b64 = base64.b64encode(f.read()).decode()
resp = client.chat.completions.create(
model="anthropic/claude-sonnet-4-5",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Turn this receipt into JSON."},
{"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{b64}"}},
],
}],
response_format={"type": "json_object"},
)
print(resp.choices[0].message.content)
Node / TypeScript
import { readFile } from "node:fs/promises";
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.alphacurve.io/v1",
apiKey: process.env.INFERENCE_API_KEY!,
});
const b64 = (await readFile("receipt.png")).toString("base64");
const resp = await client.chat.completions.create({
model: "anthropic/claude-sonnet-4-5",
messages: [{
role: "user",
content: [
{ type: "text", text: "Turn this receipt into JSON." },
{ type: "image_url", image_url: { url: `data:image/png;base64,${b64}` } },
],
}],
response_format: { type: "json_object" },
});
console.log(resp.choices[0].message.content);
Multiple images
A single message may carry several image_url parts, which is ideal for comparisons:
content = [{"type": "text", "text": "What differs between these two images?"}]
for url in ["https://example.com/a.png", "https://example.com/b.png"]:
content.append({"type": "image_url", "image_url": {"url": url}})
resp = client.chat.completions.create(
model="openai/gpt-4o",
messages=[{"role": "user", "content": content}],
)
Billing
Images are converted into input tokens. The cost depends on resolution and the detail setting. The authoritative number is always usage.prompt_tokens in the response.
To keep costs down:
- Downscale before uploading. Oversized images do not improve accuracy, only price.
- Use
detail: "low"when you only need to know what something is; reserve"high"for small text and fine detail. - Ask all your questions about an image in one message rather than re-sending it.
Notes
- Calling a model without vision support may return
400 invalid_requestupstream, or the image may simply be ignored. Checkcapabilities.visionviaGET /v1/modelsfirst. - Image URLs must be fetchable by the upstream provider. Use base64 for links that require authentication or signing.
- There is a maximum overall request size. For several high-resolution images, prefer URLs over base64.