Tool Calling 的本质是”模型给指令,程序去执行”——LLM 不直接运行函数,而是按 JSON Schema 输出一份结构化调用指令,由应用层执行后把结果回填给模型,模型再决定下一步。完整流程包含 schema 定义、意图识别、工具选择、参数生成、执行回传、五到七个步骤的循环(ReAct),每一步都依赖受约束解码保证输出合法。下文把整条链路、主流平台差异与一段 OpenAI Python 代码逐项拆开。
一、Tool Calling 与 Function Calling 的关系
OpenAI 在 2023 年 6 月推出 Function Calling API,2023 年底把 functions 参数升级为更通用的 tools 数组,并从 2024 年 6 月起在 strict: true 模式下保证 100% 符合 JSON Schema。如今”Tool Calling”已经成为行业统称——Anthropic 的 tool_use、Google Gemini 的 function calling、阿里百炼的 tool call 走的是同一套思路。模型在 SFT 阶段被喂入大量”用户问 → 选哪个工具 → 填什么参数 → 工具回返 → 整合答案”的对话样本,学会了”何时调、调哪个、怎么填”三项能力。
二、Tool Calling 的完整七步流程
从用户发问到最终答案,主流 Agent 框架(LangGraph、Strands、CrewAI)都按以下七步推进:
- 定义工具 schema:在请求里以 JSON Schema 描述工具名、描述、参数结构、必填项;
- 拼装请求:把
tools数组与messages一并发给模型; - 意图识别:模型判断”当前问句是否需要调工具”,不需要则直接生成自然语言回答;
- 工具选择 + 参数生成:需要工具时,模型按受约束解码(Constrained Decoding)生成合规 JSON;
- 解析 tool_calls:应用层读取
tool_calls[i].function.name与arguments; - 执行真实函数:本地函数或 HTTP API 拿到结果;
- 回传 observation:把结果以
tool角色消息塞回 messages,触发下一轮推理,直到finish_reason变成stop或达到最大步数。
整个循环在 ReAct 框架下被表达为 Thought → Action → Observation 的迭代。Thought 是模型的内部推理,Action 是结构化工具调用,Observation 是工具回传的实际数据。工程上有两个常被忽略的细节:其一是 tool_choice 参数可以取 auto / required / none 或指定具体工具名,让上层在”必须调””禁止调””指定调”之间做精细控制;其二是 OpenAI 从 GPT-4o 起支持 parallel tool calls,单次响应里能返回多个 tool_calls,应用层并行执行后一次性回传,能把串联任务的总延迟压到与最慢一个工具调用相近。
三、JSON Schema 设计的硬性原则
Schema 是 LLM 理解工具能力的唯一接口,描述质量直接决定调用准确率。Gorilla 与 ToolAlpaca 的研究证明,精确的 description 与 enum 约束可让参数生成准确率提升超过 30%。下面是 6 条工程原则:
| 原则 | 含义 | 反例 |
|---|---|---|
| 名称 snake_case 动词在前 | get_weather 而非 weather |
doSomething |
| description 写明边界 | 标注”何时用 / 不用” | 只说”查询天气” |
| 离散取值用 enum | 限定 ["celsius", "fahrenheit"] |
任意字符串 |
| 必填/选填要清晰 | 区分 required 数组 |
所有字段都必填 |
| 嵌套不超过 3 层 | 减少 LLM 认知负担 | 5 层对象嵌套 |
| 整型用 integer 而非 number | 防止模型填 1.5 | type: number |
工具数量超过 20 个时,description 之间的语义区分度会显著影响选择正确率。ToolLLM 的实验显示,加入”边界条件描述”(如”仅用于产品搜索,订单相关请用 getorderstatus”)能让选择错误率下降约 40%。
四、主流平台实现差异
虽然思路一致,但各家的 payload 格式并不完全兼容:
| 平台 | 工具声明位置 | 工具调用返回 | 回传角色 | 特色机制 |
|---|---|---|---|---|
| OpenAI | tools[].function |
tool_calls[].function.arguments |
role: "tool" |
parallel tool calls;strict: true |
| Anthropic Claude | tools[] |
content[].type: "tool_use" |
tool_result content block |
extended thinking;human-in-the-loop |
| Google Gemini | tools[].functionDeclarations |
functionCall 字段 |
functionResponse 字段 |
原生多模态 |
| LangChain | bind_tools([...]) |
AIMessage.tool_calls |
ToolMessage |
统一抽象 |
迁移到新平台时,工具描述文本一般可直接复用,但 payload 结构、回传角色名、并行调用规则都得按新规范重写。
五、Python 实战示例
下面给出一段最简的 OpenAI Tool Calling 流程:定义一个查天气的工具,模型自动判断调用、执行、回传,最终给用户自然语言答复。
import json
from openai import OpenAI
client = OpenAI()
# 1. 工具 schema:JSON Schema 描述
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "查询指定城市的当前天气,仅在用户问天气时使用",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "城市名,如 'Beijing'"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["city", "unit"],
"additionalProperties": False
},
"strict": True
}
}]
# 2. 第一次请求:让模型决定要不要调工具
resp = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "上海今天多少度?"}],
tools=tools,
tool_choice="auto",
)
msg = resp.choices[0].message
tool_calls = msg.tool_calls or []
# 3. 执行工具并回传 observation
messages = [{"role": "user", "content": "上海今天多少度?"}, msg]
for tc in tool_calls:
args = json.loads(tc.function.arguments)
# 真实场景:调用天气 API;这里用假数据
observation = {"city": args["city"], "temp": 28, "unit": args["unit"]}
messages.append({
"role": "tool",
"tool_call_id": tc.id,
"content": json.dumps(observation, ensure_ascii=False),
})
# 4. 第二次请求:让模型基于 observation 整合答案
final = client.chat.completions.create(model="gpt-4o", messages=messages)
print(final.choices[0].message.content)
代码块前后有几个易错点:① additionalProperties: False 必须配合 strict: True 才能保证合规;② tool_call_id 必须与第一次响应里的 id 一一对应,否则模型无法把 observation 与 call 关联;③ 如果把假数据替换为真实 API 调用的结果,记得在异常分支里返回错误 JSON,模型会自动决定重试或换工具。finish_reason 为 tool_calls 表示还要继续调用,为 stop 才算本轮收尾。
到这里,Tool Calling 从定义 schema 到回传 observation 的链路就完整了。
常见问题(FAQ)
Q1:Tool Calling 是模型自己执行函数吗?
不是,模型只输出结构化 JSON,应用层负责执行并回传结果。
Q2:JSON Schema 描述写得越详细越好吗?
描述要写边界与时机,过度冗长反而会分散模型注意力。
Q3:Tool Calling 与 ReAct 是什么关系?
ReAct 是”推理-行动”循环,Tool Calling 是循环里 Action 一步的具体实现方式。