Messages API 是 Anthropic Claude 的核心对话接口,支持 Claude 3.5、Claude 4 系列模型。该 API 设计简洁,支持多模态输入、工具调用和流式传输。
注意:本文档中的模型名称请以 Anthropic 官方文档 为准,模型版本会持续更新。
Claude API 的独特之处
Claude API 与 OpenAI 有几个关键差异,理解这些差异对于正确使用 API 至关重要:
1. System 提示词是独立参数
OpenAI 将 system 消息放在 messages 数组中,而 Claude 将其作为顶级参数。这种设计使系统提示更加突出,也便于实现 Prompt Caching。
2. max_tokens 是必需参数
OpenAI 的 max_tokens 可选,Claude 则必须指定。这是一个有意的设计决策,强制开发者考虑输出长度,避免意外的高成本。
3. 内容块设计
Claude 的消息内容使用”内容块”(content blocks)设计,每个块有明确的类型。这种设计天然支持多模态,也使得工具调用的响应更加结构化。
4. Prompt Caching
Claude 独有的功能,可以缓存重复的系统提示,在多轮对话或批量处理时节省高达 90% 的输入成本。
端点与认证
端点
POST https://api.anthropic.com/v1/messages
POST https://api.anthropic.com/v1/messages/count_tokens
/messages 是核心对话端点,/count_tokens 用于预估 Token 数量,便于成本控制。
认证
x-api-key: $ANTHROPIC_API_KEY
anthropic-version: 2023-06-01
认证方式说明:
Claude 使用自定义的 x-api-key 头而非标准的 Authorization: Bearer 格式。同时必须指定 anthropic-version 头,这是 API 版本控制的一部分,确保行为一致性。
| Header | 说明 | 是否必需 |
|---|---|---|
x-api-key | API 密钥 | 是 |
anthropic-version | API 版本 | 是 |
content-type | 固定为 application/json | 是 |
anthropic-beta | 启用 beta 功能 | 否 |
请求参数
必需参数
| 参数 | 类型 | 说明 |
|---|---|---|
model | string | 模型 ID |
messages | array | 消息数组 |
max_tokens | integer | 最大生成 token 数(必需) |
max_tokens 设置建议:
| 场景 | 建议值 | 说明 |
|---|---|---|
| 简短回答 | 256-512 | 问答、分类 |
| 一般对话 | 1024-2048 | 日常聊天 |
| 长文生成 | 4096+ | 文章、报告 |
| 代码生成 | 4096-8192 | 完整代码文件 |
可选参数
| 参数 | 类型 | 默认值 | 说明 |
|---|---|---|---|
system | string/array | - | 系统提示词 |
temperature | number | 1.0 | 随机性 0.0-1.0 |
top_p | number | - | 核采样 |
top_k | integer | - | Top-K 采样 |
stop_sequences | array | - | 停止序列 |
stream | boolean | false | 是否流式输出 |
tools | array | - | 工具定义 |
tool_choice | object | - | 工具选择策略 |
metadata | object | - | 请求元数据 |
temperature 范围差异:
注意 Claude 的 temperature 范围是 0.0-1.0,而 OpenAI 是 0-2。Claude 的 1.0 大约相当于 OpenAI 的 0.7-0.8。
消息格式
角色类型
| 角色 | 说明 |
|---|---|
user | 用户消息 |
assistant | AI 回复 |
注意:Claude 没有 system 角色,系统指令通过顶层 system 参数传递。这是与 OpenAI 最大的格式差异。
消息交替规则:
Claude 要求消息必须严格交替:user → assistant → user → assistant…。不能有连续的同角色消息。如果需要合并多条用户消息,应该在内容中合并。
基础消息
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Hello, Claude!"}
]
}
带系统指令
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"system": "You are a helpful assistant that speaks like a pirate.",
"messages": [
{"role": "user", "content": "Hello!"}
]
}
system 参数的高级用法:
system 可以是字符串,也可以是内容块数组,后者支持 Prompt Caching:
{
"system": [
{
"type": "text",
"text": "你是一个专业的技术顾问...",
"cache_control": {"type": "ephemeral"}
}
]
}
多轮对话
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "What is 2+2?"},
{"role": "assistant", "content": "2+2 equals 4."},
{"role": "user", "content": "And what is that times 3?"}
]
}
内容块类型
Claude 支持多种内容块类型:
文本块
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image:"}
]
}
图像块(Base64)
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": "/9j/4AAQSkZJRg..."
}
},
{"type": "text", "text": "What's in this image?"}
]
}
图像块(URL)
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "url",
"url": "https://example.com/image.jpg"
}
}
]
}
文档块(PDF)
{
"role": "user",
"content": [
{
"type": "document",
"source": {
"type": "base64",
"media_type": "application/pdf",
"data": "JVBERi0xLjQK..."
}
},
{"type": "text", "text": "Summarize this document"}
]
}
响应格式
Claude 的响应结构与 OpenAI 有显著不同,采用内容块数组设计。
标准响应
{
"id": "msg_01XFDUDYJgAACzvnptvVoYEL",
"type": "message",
"role": "assistant",
"content": [
{
"type": "text",
"text": "Hello! How can I help you today?"
}
],
"model": "claude-sonnet-4-20250514",
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 12,
"output_tokens": 10,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0
}
}
响应字段详解:
| 字段 | 说明 |
|---|---|
id | 消息唯一标识符,以 msg_ 开头 |
type | 固定为 message |
role | 固定为 assistant |
content | 内容块数组,可包含文本、工具调用等 |
model | 实际使用的模型版本 |
stop_reason | 停止原因 |
usage | Token 使用统计,包含缓存信息 |
content 数组的设计优势:
Claude 的 content 是数组而非字符串,这使得:
- 可以在一个响应中包含多种类型的内容
- 工具调用和文本可以混合
- 结构更清晰,便于程序解析
stop_reason 值
| 值 | 说明 | 处理建议 |
|---|---|---|
end_turn | 正常结束 | 直接使用响应 |
max_tokens | 达到 token 限制 | 考虑增加 max_tokens 或继续生成 |
stop_sequence | 遇到停止序列 | 正常情况 |
tool_use | 调用工具 | 执行工具并返回结果 |
流式传输 (SSE)
Claude 的流式传输使用更细粒度的事件类型,提供了完整的生命周期控制。
请求
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}
事件类型
| 事件 | 说明 | 包含数据 |
|---|---|---|
message_start | 消息开始 | 消息 ID、模型信息 |
content_block_start | 内容块开始 | 块索引、块类型 |
content_block_delta | 内容增量 | 文本片段 |
content_block_stop | 内容块结束 | 块索引 |
message_delta | 消息增量 | stop_reason、usage |
message_stop | 消息结束 | 无 |
事件流程图:
sequenceDiagram
participant Client
participant Claude
Claude->>Client: message_start
Claude->>Client: content_block_start (index=0)
loop 文本生成
Claude->>Client: content_block_delta
end
Claude->>Client: content_block_stop
Claude->>Client: message_delta (usage)
Claude->>Client: message_stop流式响应示例
event: message_start
data: {"type":"message_start","message":{"id":"msg_abc","type":"message","role":"assistant","content":[],"model":"claude-sonnet-4-20250514"}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"!"}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":10}}
event: message_stop
data: {"type":"message_stop"}
与 OpenAI 流式的差异:
| 方面 | OpenAI | Claude |
|---|---|---|
| 事件格式 | data: {...} | event: xxx\ndata: {...} |
| 结束标记 | data: [DONE] | event: message_stop |
| 内容字段 | delta.content | delta.text |
| Usage | 需要额外参数 | 自动包含在 message_delta |
Tool Use(工具调用)
Claude 的工具调用与 OpenAI 类似,但有一些细节差异。
工作流程
sequenceDiagram
participant User
participant App
participant Claude
participant Tool
User->>App: 东京天气怎么样?
App->>Claude: 消息 + 工具定义
Claude->>App: tool_use 内容块
App->>Tool: 调用天气 API
Tool->>App: 返回天气数据
App->>Claude: tool_result 消息
Claude->>App: 最终回复
App->>User: 东京今天晴,22°C定义工具
注意 Claude 使用 input_schema 而非 OpenAI 的 parameters:
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"tools": [
{
"name": "get_weather",
"description": "Get the current weather for a location",
"input_schema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City name, e.g. San Francisco, CA"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit"
}
},
"required": ["location"]
}
}
],
"messages": [
{"role": "user", "content": "What's the weather in Tokyo?"}
]
}
与 OpenAI 工具定义的差异:
| 方面 | OpenAI | Claude |
|---|---|---|
| 参数字段 | parameters | input_schema |
| 外层包装 | {"type": "function", "function": {...}} | 直接定义 |
| 工具 ID 前缀 | call_ | toolu_ |
工具调用响应
当 Claude 决定调用工具时,响应的 content 数组中会包含 tool_use 类型的内容块:
{
"id": "msg_abc",
"type": "message",
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "toolu_01A09q90qw90lq917835lgs",
"name": "get_weather",
"input": {"location": "Tokyo", "unit": "celsius"}
}
],
"stop_reason": "tool_use"
}
注意事项:
input已经是解析好的对象,不是 JSON 字符串(与 OpenAI 不同)id以toolu_开头stop_reason为tool_use
返回工具结果
工具结果通过 tool_result 内容块返回,注意它放在 user 角色的消息中:
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "What's the weather in Tokyo?"},
{
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "toolu_01A09q90qw90lq917835lgs",
"name": "get_weather",
"input": {"location": "Tokyo", "unit": "celsius"}
}
]
},
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_01A09q90qw90lq917835lgs",
"content": "{\"temperature\": 22, \"condition\": \"sunny\"}"
}
]
}
]
}
与 OpenAI 的差异:
| 方面 | OpenAI | Claude |
|---|---|---|
| 结果角色 | tool | user |
| ID 字段 | tool_call_id | tool_use_id |
| 内容类型 | 直接是 content | tool_result 块 |
tool_choice 选项
// 自动选择(默认)
{"type": "auto"}
// 必须使用任意工具
{"type": "any"}
// 必须使用指定工具
{"type": "tool", "name": "get_weather"}
Prompt Caching(提示词缓存)
Prompt Caching 是 Claude 的独特功能,可以缓存重复的系统提示,大幅降低成本。
工作原理
flowchart LR
A[首次请求] --> B[创建缓存]
B --> C[后续请求]
C --> D{缓存命中?}
D -->|是| E[使用缓存 -90% 成本]
D -->|否| F[重新计算]使用方法
在需要缓存的内容块上添加 cache_control:
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"system": [
{
"type": "text",
"text": "You are a helpful assistant with extensive knowledge...",
"cache_control": {"type": "ephemeral"}
}
],
"messages": [
{"role": "user", "content": "Hello"}
]
}
缓存统计
响应的 usage 中会包含缓存相关统计:
{
"usage": {
"input_tokens": 12,
"output_tokens": 10,
"cache_creation_input_tokens": 1500, // 首次创建缓存
"cache_read_input_tokens": 0 // 后续命中缓存
}
}
成本节省
| 场景 | 无缓存 | 有缓存 | 节省 |
|---|---|---|---|
| 首次请求 | 100% | 125%(创建缓存) | -25% |
| 后续请求 | 100% | 10% | 90% |
| 10 次请求平均 | 100% | 21.5% | 78.5% |
适用场景:
- 长系统提示词
- 多轮对话(缓存历史)
- 批量处理相同模板
- RAG 场景(缓存检索结果)
缓存 TTL:
ephemeral: 5 分钟- 付费用户可获得更长 TTL
完整示例
cURL
curl https://api.anthropic.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Hello, Claude!"}
]
}'
Python
import anthropic
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[
{"role": "user", "content": "Hello, Claude!"}
]
)
print(message.content[0].text)
Python 流式
import anthropic
client = anthropic.Anthropic()
with client.messages.stream(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": "Write a haiku"}]
) as stream:
for text in stream.text_stream:
print(text, end="")
Python 多模态
import anthropic
import base64
client = anthropic.Anthropic()
# 读取图片
with open("image.jpg", "rb") as f:
image_data = base64.standard_b64encode(f.read()).decode("utf-8")
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": image_data
}
},
{"type": "text", "text": "Describe this image"}
]
}
]
)
print(message.content[0].text)
Node.js
import Anthropic from '@anthropic-ai/sdk';
const anthropic = new Anthropic();
const message = await anthropic.messages.create({
model: 'claude-sonnet-4-20250514',
max_tokens: 1024,
messages: [
{ role: 'user', content: 'Hello, Claude!' }
]
});
console.log(message.content[0].text);
错误处理
常见错误码
| 状态码 | 说明 |
|---|---|
| 400 | 请求格式错误 |
| 401 | API Key 无效 |
| 403 | 无权限 |
| 404 | 资源不存在 |
| 429 | 速率限制 |
| 500 | 服务器错误 |
| 529 | API 过载 |
错误响应格式
{
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "max_tokens: Field required"
}
}
模型列表
| 模型 | 上下文窗口 | 说明 |
|---|---|---|
| claude-sonnet-4-20250514 | 200K | Claude 4 平衡版本 |
| claude-3-5-sonnet-20241022 | 200K | Claude 3.5 主力模型 |
| claude-3-5-haiku-20241022 | 200K | 快速经济型 |
| claude-3-opus-20240229 | 200K | Claude 3 旗舰(推理强) |
模型名称格式:
claude-{版本}-{变体}-{日期},请查阅官方文档获取最新可用模型。
与 OpenAI 的主要差异
| 特性 | Claude | OpenAI |
|---|---|---|
| 系统指令 | 顶层 system 参数 | messages 中的 system 角色 |
| max_tokens | 必需参数 | 可选参数 |
| 响应结构 | content[0].text | choices[0].message.content |
| 工具定义 | input_schema | parameters |
| 工具结果 | tool_result 在 user 消息中 | tool 角色消息 |
| 认证头 | x-api-key | Authorization: Bearer |
最佳实践
1. 利用 Prompt Caching 降低成本
对于重复使用的系统提示,启用缓存可节省高达 90% 的输入成本:
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
system=[{
"type": "text",
"text": "你是一个专业的代码审查助手...(长文本)",
"cache_control": {"type": "ephemeral"}
}],
messages=[{"role": "user", "content": "审查这段代码"}]
)
2. 处理长文档
Claude 支持 200K 上下文,适合处理长文档:
# 直接上传 PDF
with open("document.pdf", "rb") as f:
pdf_data = base64.b64encode(f.read()).decode()
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=4096,
messages=[{
"role": "user",
"content": [
{"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": pdf_data}},
{"type": "text", "text": "总结这份文档的要点"}
]
}]
)
3. 结构化输出
虽然 Claude 没有原生 JSON Mode,但可以通过提示词实现:
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
system="你必须以 JSON 格式回复,不要包含任何其他文本。",
messages=[{"role": "user", "content": "列出 3 种编程语言及其特点"}]
)
4. 错误处理
from anthropic import APIError, RateLimitError, APIConnectionError
try:
message = client.messages.create(...)
except RateLimitError:
print("请求过于频繁,请稍后重试")
except APIConnectionError:
print("网络连接失败")
except APIError as e:
print(f"API 错误: {e.message}")
相关指南
官方文档
- API Reference: https://docs.anthropic.com/en/api/messages
- Python SDK: https://github.com/anthropics/anthropic-sdk-python
- Tool Use Guide: https://docs.anthropic.com/en/docs/build-with-claude/tool-use
- Vision Guide: https://docs.anthropic.com/en/docs/build-with-claude/vision
- Prompt Caching: https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching