常见问题与排错
大模型 API 开发中的常见问题、错误排查和解决方案。
排错思路
遇到问题时,按以下顺序排查:
flowchart TD
A[遇到错误] --> B{有错误码?}
B -->|是| C[查错误码表]
B -->|否| D{有响应?}
D -->|是| E[检查响应内容]
D -->|否| F[检查网络/超时]
C --> G[针对性解决]
E --> G
F --> G通用排查步骤:
- 查看完整错误信息
- 检查 API Key 和权限
- 验证请求格式
- 测试网络连通性
- 查看官方状态页
认证问题
API Key 无效
Error: Invalid API key
Error: Incorrect API key provided
常见原因和解决方案:
| 原因 | 解决 |
|---|---|
| Key 复制不完整 | 重新复制,确保完整 |
| Key 已过期/撤销 | 重新生成新 Key |
| 环境变量未设置 | 检查 OPENAI_API_KEY 等 |
| Key 前后有空格 | key.strip() 去除空白 |
| 使用了错误的 Key | 确认 Key 对应的服务商 |
# 诊断脚本
import os
def check_api_key():
key = os.getenv("OPENAI_API_KEY", "")
if not key:
print("❌ 环境变量未设置")
return
print(f"✓ Key 长度: {len(key)}")
print(f"✓ Key 前缀: {key[:8]}...")
# 检查常见问题
if key != key.strip():
print("⚠️ Key 包含空白字符")
if key.startswith("sk-"):
print("✓ OpenAI Key 格式正确")
elif key.startswith("sk-ant-"):
print("✓ Claude Key 格式正确")
check_api_key()
权限不足
Error: You don't have access to this model
Error: Permission denied
| 原因 | 解决 |
|---|---|
| 账户等级不够 | 升级账户或充值 |
| 模型需要申请 | 申请 GPT-4 等模型访问权限 |
| 区域限制 | 使用代理或切换区域 |
| 组织权限 | 检查组织设置 |
请求问题
超时
Error: Request timed out
Error: Connection timed out
超时通常发生在:
- 网络不稳定
- 请求内容过长
- 服务器负载高
from openai import OpenAI
# 方案1:增加超时时间
client = OpenAI(timeout=120.0) # 默认是 60 秒
# 方案2:使用流式输出(推荐)
# 流式输出可以更快获得首个 token,避免长时间等待
for chunk in client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "长文本..."}],
stream=True
):
content = chunk.choices[0].delta.content
if content:
print(content, end="", flush=True)
速率限制
Error: Rate limit exceeded
Error: Too many requests
速率限制类型:
| 类型 | 说明 | 解决 |
|---|---|---|
| RPM | 每分钟请求数 | 降低并发 |
| TPM | 每分钟 Token 数 | 减少单次 Token |
| RPD | 每天请求数 | 升级配额 |
import time
from tenacity import retry, wait_exponential, stop_after_attempt
# 指数退避重试
@retry(
wait=wait_exponential(multiplier=1, min=1, max=60),
stop=stop_after_attempt(5)
)
def call_api(messages):
return client.chat.completions.create(
model="gpt-4o-mini",
messages=messages
)
# 简单的速率控制
class RateLimiter:
def __init__(self, calls_per_minute=60):
self.interval = 60 / calls_per_minute
self.last_call = 0
def wait(self):
elapsed = time.time() - self.last_call
if elapsed < self.interval:
time.sleep(self.interval - elapsed)
self.last_call = time.time()
limiter = RateLimiter(calls_per_minute=50)
for msg in messages:
limiter.wait()
response = call_api(msg)
请求体过大
Error: Request body too large
Error: Maximum context length exceeded
| 原因 | 解决 |
|---|---|
| 消息太长 | 截断或摘要压缩 |
| 图片太大 | 压缩图片或降低分辨率 |
| 历史太多 | 限制历史条数 |
def trim_messages(messages, max_messages=20):
"""保留 system 消息和最近的对话"""
system = [m for m in messages if m["role"] == "system"]
others = [m for m in messages if m["role"] != "system"]
return system + others[-max_messages:]
# 更智能的截断:基于 Token 估算
def trim_by_tokens(messages, max_tokens=4000):
"""基于 Token 数量截断"""
system = [m for m in messages if m["role"] == "system"]
others = [m for m in messages if m["role"] != "system"]
result = system.copy()
total = sum(len(m["content"]) // 4 for m in system) # 粗略估算
for m in reversed(others):
tokens = len(m["content"]) // 4
if total + tokens > max_tokens:
break
result.insert(len(system), m)
total += tokens
return result
响应问题
空响应
响应为空或 choices 数组为空:
response = client.chat.completions.create(...)
# 检查响应
if not response.choices:
print("❌ 空响应")
elif not response.choices[0].message.content:
print("❌ 内容为空")
原因和解决:
| 原因 | 解决 |
|---|---|
| 内容被安全过滤 | 检查 finish_reason |
| max_tokens 太小 | 增加 max_tokens |
| 模型拒绝回答 | 修改 prompt |
| 触发内容策略 | 调整敏感内容 |
# 检查完成原因
choice = response.choices[0]
finish_reason = choice.finish_reason
if finish_reason == "content_filter":
print("⚠️ 内容被安全过滤")
elif finish_reason == "length":
print("⚠️ 达到长度限制,输出被截断")
elif finish_reason == "stop":
print("✓ 正常完成")
elif finish_reason == "tool_calls":
print("✓ 需要调用工具")
JSON 格式错误
LLM 返回的 JSON 可能格式不正确:
import json
import re
content = response.choices[0].message.content
def safe_parse_json(content):
"""安全解析 JSON,处理常见问题"""
# 1. 直接尝试解析
try:
return json.loads(content)
except json.JSONDecodeError:
pass
# 2. 尝试提取 JSON 块
# 处理 ```json ... ``` 格式
match = re.search(r'```(?:json)?\s*([\s\S]*?)```', content)
if match:
try:
return json.loads(match.group(1))
except:
pass
# 3. 尝试提取 {...} 或 [...]
match = re.search(r'(\{[\s\S]*\}|\[[\s\S]*\])', content)
if match:
try:
return json.loads(match.group(1))
except:
pass
raise ValueError(f"无法解析 JSON: {content[:100]}...")
# 使用
try:
data = safe_parse_json(content)
except ValueError as e:
print(f"解析失败: {e}")
更好的方案:使用 JSON Mode
# OpenAI JSON Mode
response = client.chat.completions.create(
model="gpt-4o",
messages=[...],
response_format={"type": "json_object"}
)
# 返回保证是有效 JSON
乱码/编码问题
# 确保 UTF-8 编码
import requests
response = requests.post(url, json=data)
response.encoding = 'utf-8'
content = response.json()
# 处理特殊字符
text = content.encode('utf-8').decode('utf-8')
流式输出问题
流式解析错误
# 正确处理流式响应
for chunk in response:
# 检查是否有内容
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
# 完整的流式处理
def stream_response(response):
full_content = ""
for chunk in response:
if not chunk.choices:
continue
delta = chunk.choices[0].delta
if delta.content:
full_content += delta.content
yield delta.content
return full_content
SSE 连接断开
import httpx
# 增加超时和重试
async def stream_with_retry(url, data, max_retries=3):
for attempt in range(max_retries):
try:
async with httpx.AsyncClient(timeout=None) as client:
async with client.stream("POST", url, json=data) as response:
async for line in response.aiter_lines():
if line.startswith("data: "):
yield line[6:]
break
except httpx.ReadTimeout:
if attempt < max_retries - 1:
print(f"重试 {attempt + 1}/{max_retries}")
else:
raise
流式输出不完整
# 确保收集完整响应
full_response = ""
finish_reason = None
for chunk in response:
if chunk.choices:
delta = chunk.choices[0].delta
if delta.content:
full_response += delta.content
if chunk.choices[0].finish_reason:
finish_reason = chunk.choices[0].finish_reason
if finish_reason != "stop":
print(f"⚠️ 异常结束: {finish_reason}")
工具调用问题
工具未被调用
LLM 没有调用你定义的工具:
| 原因 | 解决 |
|---|---|
| 工具描述不清晰 | 优化 description |
| 参数定义错误 | 检查 JSON Schema |
| 模型不支持 | 使用支持工具的模型 |
| Prompt 不明确 | 明确告诉模型可以使用工具 |
# 好的工具定义示例
tools = [{
"type": "function",
"function": {
"name": "get_weather",
# description 要清晰说明何时使用
"description": "获取指定城市的当前天气信息。当用户询问天气、温度、是否下雨等问题时使用此工具。",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
# 参数描述也要清晰
"description": "城市名称,如:北京、上海、广州"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "温度单位,默认摄氏度"
}
},
"required": ["city"]
}
}
}]
参数解析错误
import json
# 安全解析工具参数
def parse_tool_args(tool_call):
try:
return json.loads(tool_call.function.arguments)
except json.JSONDecodeError as e:
print(f"参数解析失败: {e}")
print(f"原始参数: {tool_call.function.arguments}")
return {}
# 使用
tool_call = response.choices[0].message.tool_calls[0]
args = parse_tool_args(tool_call)
工具调用循环
Agent 陷入无限工具调用循环:
# 限制最大调用次数
MAX_TOOL_CALLS = 10
def run_agent(messages, tools):
for i in range(MAX_TOOL_CALLS):
response = client.chat.completions.create(
model="gpt-4o",
messages=messages,
tools=tools
)
if response.choices[0].finish_reason == "stop":
return response.choices[0].message.content
# 处理工具调用...
raise Exception("达到最大工具调用次数")
网络问题
代理设置
import httpx
from openai import OpenAI
# 方法1:通过 httpx 设置代理
client = OpenAI(
http_client=httpx.Client(
proxies="http://127.0.0.1:7890"
)
)
# 方法2:环境变量
# export HTTP_PROXY=http://127.0.0.1:7890
# export HTTPS_PROXY=http://127.0.0.1:7890
# export ALL_PROXY=socks5://127.0.0.1:7890
# 测试代理是否工作
curl -x http://127.0.0.1:7890 https://api.openai.com/v1/models \
-H "Authorization: Bearer $OPENAI_API_KEY"
SSL 证书错误
Error: SSL certificate verify failed
import httpx
# 临时禁用验证(仅用于调试,不推荐生产使用)
client = OpenAI(
http_client=httpx.Client(verify=False)
)
# 更好的方案:指定证书
client = OpenAI(
http_client=httpx.Client(verify="/path/to/cert.pem")
)
DNS 解析失败
# 使用自定义 DNS 或直接 IP
import httpx
# 指定 base_url
client = OpenAI(
base_url="https://api.openai.com/v1", # 或使用 IP
http_client=httpx.Client(
transport=httpx.HTTPTransport(
retries=3
)
)
)
调试技巧
1. 启用日志
import logging
# 设置日志级别
logging.basicConfig(level=logging.DEBUG)
# OpenAI SDK 专用日志
import openai
openai.log = "debug"
2. 打印请求/响应
import httpx
def log_request(request):
print(f"→ {request.method} {request.url}")
print(f" Headers: {dict(request.headers)}")
if request.content:
print(f" Body: {request.content[:500]}...")
def log_response(response):
print(f"← {response.status_code}")
client = OpenAI(
http_client=httpx.Client(
event_hooks={
"request": [log_request],
"response": [log_response]
}
)
)
3. 使用 curl 测试
# 基础测试
curl https://api.openai.com/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello"}]
}'
# 带详细输出
curl -v https://api.openai.com/v1/models \
-H "Authorization: Bearer $OPENAI_API_KEY"
# 测试流式输出
curl https://api.openai.com/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Count to 5"}],
"stream": true
}'
4. 最小复现
# 创建最小测试用例
from openai import OpenAI
client = OpenAI()
# 最简单的请求
try:
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hi"}],
max_tokens=10
)
print("✓ 基础请求成功")
print(response.choices[0].message.content)
except Exception as e:
print(f"✗ 错误: {type(e).__name__}: {e}")
错误码速查
| 错误码 | 含义 | 常见原因 | 解决方案 |
|---|---|---|---|
| 400 | Bad Request | 请求格式错误 | 检查参数格式 |
| 401 | Unauthorized | API Key 无效 | 检查 Key |
| 403 | Forbidden | 权限不足 | 检查账户权限 |
| 404 | Not Found | 模型/端点不存在 | 检查模型名称 |
| 422 | Unprocessable | 参数验证失败 | 检查参数值 |
| 429 | Too Many Requests | 速率限制 | 降低频率/重试 |
| 500 | Internal Error | 服务器错误 | 稍后重试 |
| 502 | Bad Gateway | 网关错误 | 稍后重试 |
| 503 | Service Unavailable | 服务不可用 | 稍后重试 |
| 504 | Gateway Timeout | 网关超时 | 增加超时/重试 |
各厂商状态页:
| 厂商 | 状态页 |
|---|---|
| OpenAI | status.openai.com |
| Anthropic | status.anthropic.com |
| status.cloud.google.com |