Skip to content

Vision (image input)

Pass image URLs or base64 data through image_url content parts so models can read images.

Vision (image input)

Vision-capable models can read images. Change content from a string to an array of content parts, mixing text and image_url entries.

Content part format

{
  "role": "user",
  "content": [
    { "type": "text", "text": "What is in this image?" },
    {
      "type": "image_url",
      "image_url": {
        "url": "https://example.com/photo.jpg",
        "detail": "auto"
      }
    }
  ]
}
FieldDescription
image_url.urlA publicly reachable HTTPS URL, or a data: base64 URI.
image_url.detailauto (default), low, high. low uses fewer tokens and costs less; high preserves detail. Some upstream models ignore this field.

Supported formats: PNG, JPEG, WebP, and non-animated GIF.

Passing a URL

curl

curl https://api.alphacurve.io/v1/chat/completions \
  -H "Authorization: Bearer $INFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-4o",
    "messages": [{
      "role": "user",
      "content": [
        { "type": "text", "text": "What is in this image?" },
        { "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } }
      ]
    }]
  }'

Python

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.alphacurve.io/v1",
    api_key=os.environ["INFERENCE_API_KEY"],
)

resp = client.chat.completions.create(
    model="openai/gpt-4o",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_url",
             "image_url": {"url": "https://example.com/photo.jpg"}},
        ],
    }],
)

print(resp.choices[0].message.content)

Node / TypeScript

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.alphacurve.io/v1",
  apiKey: process.env.INFERENCE_API_KEY!,
});

const resp = await client.chat.completions.create({
  model: "google/gemini-2.5-pro",
  messages: [{
    role: "user",
    content: [
      { type: "text", text: "What is in this image?" },
      { type: "image_url", image_url: { url: "https://example.com/photo.jpg" } },
    ],
  }],
});

console.log(resp.choices[0].message.content);

Passing a local file (base64)

When the image is not on the public internet, use a data: URI.

Python

import base64
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.alphacurve.io/v1",
    api_key=os.environ["INFERENCE_API_KEY"],
)

with open("receipt.png", "rb") as f:
    b64 = base64.b64encode(f.read()).decode()

resp = client.chat.completions.create(
    model="anthropic/claude-sonnet-4-5",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Turn this receipt into JSON."},
            {"type": "image_url",
             "image_url": {"url": f"data:image/png;base64,{b64}"}},
        ],
    }],
    response_format={"type": "json_object"},
)

print(resp.choices[0].message.content)

Node / TypeScript

import { readFile } from "node:fs/promises";
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.alphacurve.io/v1",
  apiKey: process.env.INFERENCE_API_KEY!,
});

const b64 = (await readFile("receipt.png")).toString("base64");

const resp = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-4-5",
  messages: [{
    role: "user",
    content: [
      { type: "text", text: "Turn this receipt into JSON." },
      { type: "image_url", image_url: { url: `data:image/png;base64,${b64}` } },
    ],
  }],
  response_format: { type: "json_object" },
});

console.log(resp.choices[0].message.content);

Multiple images

A single message may carry several image_url parts, which is ideal for comparisons:

content = [{"type": "text", "text": "What differs between these two images?"}]
for url in ["https://example.com/a.png", "https://example.com/b.png"]:
    content.append({"type": "image_url", "image_url": {"url": url}})

resp = client.chat.completions.create(
    model="openai/gpt-4o",
    messages=[{"role": "user", "content": content}],
)

Billing

Images are converted into input tokens. The cost depends on resolution and the detail setting. The authoritative number is always usage.prompt_tokens in the response.

To keep costs down:

  • Downscale before uploading. Oversized images do not improve accuracy, only price.
  • Use detail: "low" when you only need to know what something is; reserve "high" for small text and fine detail.
  • Ask all your questions about an image in one message rather than re-sending it.

Notes

  • Calling a model without vision support may return 400 invalid_request upstream, or the image may simply be ignored. Check capabilities.vision via GET /v1/models first.
  • Image URLs must be fetchable by the upstream provider. Use base64 for links that require authentication or signing.
  • There is a maximum overall request size. For several high-resolution images, prefer URLs over base64.