Skip to content

Tools & function calling

Let the model call your functions: tool definitions, tool_choice, the full three-step loop and streaming.

Tools & function calling

Tool calling gives a model access to external capabilities. The model does not execute anything itself — it tells you "call this function with these arguments". You run it, hand the result back, and the model produces the final answer.

Inference normalises every provider's tool interface into the OpenAI format, so the same tool definitions work across OpenAI, Anthropic and Gemini models.

Tool definition

{
  "type": "function",
  "function": {
    "name": "get_weather",
    "description": "Get the current weather for a city.",
    "parameters": {
      "type": "object",
      "properties": {
        "city": { "type": "string", "description": "City name, e.g. 'Taipei'" },
        "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
      },
      "required": ["city"]
    }
  }
}

The clearer the description, the better the model's judgement. parameters uses JSON Schema.

tool_choice

ValueBehaviour
"auto"The model decides whether to call a tool. Default when tools is present.
"none"No tool calls; text only.
"required"Force at least one tool call.
{"type": "function", "function": {"name": "get_weather"}}Force a specific tool.

The three-step loop

Step 1 — send the request with tools

curl https://api.alphacurve.io/v1/chat/completions \
  -H "Authorization: Bearer $INFERENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-4o",
    "messages": [{ "role": "user", "content": "What is the weather in Taipei?" }],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": { "city": { "type": "string" } },
          "required": ["city"]
        }
      }
    }],
    "tool_choice": "auto"
  }'

Step 2 — the model returns tool_calls

{
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_a1b2c3",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"city\":\"Taipei\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ]
}

Note that arguments is a JSON string, not an object. Parse it yourself.

Step 3 — send the result back

The tool message must carry a tool_call_id matching the id above.

{
  "messages": [
    { "role": "user", "content": "What is the weather in Taipei?" },
    {
      "role": "assistant",
      "content": null,
      "tool_calls": [{ "id": "call_a1b2c3", "type": "function",
        "function": { "name": "get_weather", "arguments": "{\"city\":\"Taipei\"}" } }]
    },
    {
      "role": "tool",
      "tool_call_id": "call_a1b2c3",
      "content": "{\"temp_c\": 31, \"condition\": \"cloudy\"}"
    }
  ]
}

Full example (Python)

import json
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.alphacurve.io/v1",
    api_key=os.environ["INFERENCE_API_KEY"],
)

TOOLS = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

def get_weather(city: str) -> dict:
    return {"city": city, "temp_c": 31, "condition": "cloudy"}

messages = [{"role": "user", "content": "What is the weather in Taipei?"}]

while True:
    resp = client.chat.completions.create(
        model="openai/gpt-4o",
        messages=messages,
        tools=TOOLS,
        tool_choice="auto",
    )
    msg = resp.choices[0].message
    messages.append(msg.model_dump(exclude_none=True))

    if not msg.tool_calls:
        print(msg.content)
        break

    for call in msg.tool_calls:
        args = json.loads(call.function.arguments)
        result = get_weather(**args)
        messages.append({
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result),
        })

Full example (Node / TypeScript)

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.alphacurve.io/v1",
  apiKey: process.env.INFERENCE_API_KEY!,
});

const tools: OpenAI.Chat.ChatCompletionTool[] = [{
  type: "function",
  function: {
    name: "get_weather",
    description: "Get the current weather for a city.",
    parameters: {
      type: "object",
      properties: { city: { type: "string" } },
      required: ["city"],
    },
  },
}];

function getWeather(city: string) {
  return { city, temp_c: 31, condition: "cloudy" };
}

const messages: OpenAI.Chat.ChatCompletionMessageParam[] = [
  { role: "user", content: "What is the weather in Taipei?" },
];

for (;;) {
  const resp = await client.chat.completions.create({
    model: "openai/gpt-4o",
    messages,
    tools,
    tool_choice: "auto",
  });

  const msg = resp.choices[0].message;
  messages.push(msg);

  if (!msg.tool_calls?.length) {
    console.log(msg.content);
    break;
  }

  for (const call of msg.tool_calls) {
    const args = JSON.parse(call.function.arguments);
    messages.push({
      role: "tool",
      tool_call_id: call.id,
      content: JSON.stringify(getWeather(args.city)),
    });
  }
}

Tool calls while streaming

When streaming, arguments arrives in fragments. Accumulate them by index:

acc = {}

for chunk in client.chat.completions.create(
    model="openai/gpt-4o",
    messages=[{"role": "user", "content": "Weather in Taipei?"}],
    tools=TOOLS,
    stream=True,
):
    for tc in (chunk.choices[0].delta.tool_calls or []):
        slot = acc.setdefault(tc.index, {"id": "", "name": "", "arguments": ""})
        if tc.id:
            slot["id"] = tc.id
        if tc.function and tc.function.name:
            slot["name"] = tc.function.name
        if tc.function and tc.function.arguments:
            slot["arguments"] += tc.function.arguments

print(acc)  # only json.loads once arguments is complete

Practical guidance

  • Always validate arguments. Model-generated arguments is untrusted input. Validate against your schema before executing.
  • Bound the loop. Cap the number of rounds (say 8) so a runaway tool loop cannot drain your balance.
  • Keep the tool list small. Definitions consume input tokens, and too many options degrade selection accuracy.
  • Return errors too. When a tool fails, send the error text back as the tool message content — models usually recover on their own.
  • Check model support via GET /v1/models and the capabilities.tools flag.