全部笔记All notes

大模型 API 常见问题与排错

阅读 7m 40s7m 40s read

常见问题与排错

大模型 API 开发中的常见问题、错误排查和解决方案。


排错思路

遇到问题时,按以下顺序排查:

flowchart TD
    A[遇到错误] --> B{有错误码?}
    B -->|是| C[查错误码表]
    B -->|否| D{有响应?}
    D -->|是| E[检查响应内容]
    D -->|否| F[检查网络/超时]
    C --> G[针对性解决]
    E --> G
    F --> G

通用排查步骤:

  1. 查看完整错误信息
  2. 检查 API Key 和权限
  3. 验证请求格式
  4. 测试网络连通性
  5. 查看官方状态页

认证问题

API Key 无效

Error: Invalid API key
Error: Incorrect API key provided

常见原因和解决方案:

原因解决
Key 复制不完整重新复制,确保完整
Key 已过期/撤销重新生成新 Key
环境变量未设置检查 OPENAI_API_KEY 等
Key 前后有空格key.strip() 去除空白
使用了错误的 Key确认 Key 对应的服务商
# 诊断脚本
import os

def check_api_key():
    key = os.getenv("OPENAI_API_KEY", "")
    
    if not key:
        print("❌ 环境变量未设置")
        return
    
    print(f"✓ Key 长度: {len(key)}")
    print(f"✓ Key 前缀: {key[:8]}...")
    
    # 检查常见问题
    if key != key.strip():
        print("⚠️ Key 包含空白字符")
    if key.startswith("sk-"):
        print("✓ OpenAI Key 格式正确")
    elif key.startswith("sk-ant-"):
        print("✓ Claude Key 格式正确")

check_api_key()

权限不足

Error: You don't have access to this model
Error: Permission denied
原因解决
账户等级不够升级账户或充值
模型需要申请申请 GPT-4 等模型访问权限
区域限制使用代理或切换区域
组织权限检查组织设置

请求问题

超时

Error: Request timed out
Error: Connection timed out

超时通常发生在:

  • 网络不稳定
  • 请求内容过长
  • 服务器负载高
from openai import OpenAI

# 方案1:增加超时时间
client = OpenAI(timeout=120.0)  # 默认是 60 秒

# 方案2:使用流式输出(推荐)
# 流式输出可以更快获得首个 token,避免长时间等待
for chunk in client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "长文本..."}],
    stream=True
):
    content = chunk.choices[0].delta.content
    if content:
        print(content, end="", flush=True)

速率限制

Error: Rate limit exceeded
Error: Too many requests

速率限制类型:

类型说明解决
RPM每分钟请求数降低并发
TPM每分钟 Token 数减少单次 Token
RPD每天请求数升级配额
import time
from tenacity import retry, wait_exponential, stop_after_attempt

# 指数退避重试
@retry(
    wait=wait_exponential(multiplier=1, min=1, max=60),
    stop=stop_after_attempt(5)
)
def call_api(messages):
    return client.chat.completions.create(
        model="gpt-4o-mini",
        messages=messages
    )

# 简单的速率控制
class RateLimiter:
    def __init__(self, calls_per_minute=60):
        self.interval = 60 / calls_per_minute
        self.last_call = 0
    
    def wait(self):
        elapsed = time.time() - self.last_call
        if elapsed < self.interval:
            time.sleep(self.interval - elapsed)
        self.last_call = time.time()

limiter = RateLimiter(calls_per_minute=50)
for msg in messages:
    limiter.wait()
    response = call_api(msg)

请求体过大

Error: Request body too large
Error: Maximum context length exceeded
原因解决
消息太长截断或摘要压缩
图片太大压缩图片或降低分辨率
历史太多限制历史条数
def trim_messages(messages, max_messages=20):
    """保留 system 消息和最近的对话"""
    system = [m for m in messages if m["role"] == "system"]
    others = [m for m in messages if m["role"] != "system"]
    return system + others[-max_messages:]

# 更智能的截断:基于 Token 估算
def trim_by_tokens(messages, max_tokens=4000):
    """基于 Token 数量截断"""
    system = [m for m in messages if m["role"] == "system"]
    others = [m for m in messages if m["role"] != "system"]
    
    result = system.copy()
    total = sum(len(m["content"]) // 4 for m in system)  # 粗略估算
    
    for m in reversed(others):
        tokens = len(m["content"]) // 4
        if total + tokens > max_tokens:
            break
        result.insert(len(system), m)
        total += tokens
    
    return result

响应问题

空响应

响应为空或 choices 数组为空:

response = client.chat.completions.create(...)

# 检查响应
if not response.choices:
    print("❌ 空响应")
elif not response.choices[0].message.content:
    print("❌ 内容为空")

原因和解决:

原因解决
内容被安全过滤检查 finish_reason
max_tokens 太小增加 max_tokens
模型拒绝回答修改 prompt
触发内容策略调整敏感内容
# 检查完成原因
choice = response.choices[0]
finish_reason = choice.finish_reason

if finish_reason == "content_filter":
    print("⚠️ 内容被安全过滤")
elif finish_reason == "length":
    print("⚠️ 达到长度限制,输出被截断")
elif finish_reason == "stop":
    print("✓ 正常完成")
elif finish_reason == "tool_calls":
    print("✓ 需要调用工具")

JSON 格式错误

LLM 返回的 JSON 可能格式不正确:

import json
import re

content = response.choices[0].message.content

def safe_parse_json(content):
    """安全解析 JSON,处理常见问题"""
    # 1. 直接尝试解析
    try:
        return json.loads(content)
    except json.JSONDecodeError:
        pass
    
    # 2. 尝试提取 JSON 块
    # 处理 ```json ... ``` 格式
    match = re.search(r'```(?:json)?\s*([\s\S]*?)```', content)
    if match:
        try:
            return json.loads(match.group(1))
        except:
            pass
    
    # 3. 尝试提取 {...} 或 [...]
    match = re.search(r'(\{[\s\S]*\}|\[[\s\S]*\])', content)
    if match:
        try:
            return json.loads(match.group(1))
        except:
            pass
    
    raise ValueError(f"无法解析 JSON: {content[:100]}...")

# 使用
try:
    data = safe_parse_json(content)
except ValueError as e:
    print(f"解析失败: {e}")

更好的方案:使用 JSON Mode

# OpenAI JSON Mode
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[...],
    response_format={"type": "json_object"}
)
# 返回保证是有效 JSON

乱码/编码问题

# 确保 UTF-8 编码
import requests

response = requests.post(url, json=data)
response.encoding = 'utf-8'
content = response.json()

# 处理特殊字符
text = content.encode('utf-8').decode('utf-8')

流式输出问题

流式解析错误

# 正确处理流式响应
for chunk in response:
    # 检查是否有内容
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

# 完整的流式处理
def stream_response(response):
    full_content = ""
    for chunk in response:
        if not chunk.choices:
            continue
        delta = chunk.choices[0].delta
        if delta.content:
            full_content += delta.content
            yield delta.content
    return full_content

SSE 连接断开

import httpx

# 增加超时和重试
async def stream_with_retry(url, data, max_retries=3):
    for attempt in range(max_retries):
        try:
            async with httpx.AsyncClient(timeout=None) as client:
                async with client.stream("POST", url, json=data) as response:
                    async for line in response.aiter_lines():
                        if line.startswith("data: "):
                            yield line[6:]
            break
        except httpx.ReadTimeout:
            if attempt < max_retries - 1:
                print(f"重试 {attempt + 1}/{max_retries}")
            else:
                raise

流式输出不完整

# 确保收集完整响应
full_response = ""
finish_reason = None

for chunk in response:
    if chunk.choices:
        delta = chunk.choices[0].delta
        if delta.content:
            full_response += delta.content
        if chunk.choices[0].finish_reason:
            finish_reason = chunk.choices[0].finish_reason

if finish_reason != "stop":
    print(f"⚠️ 异常结束: {finish_reason}")

工具调用问题

工具未被调用

LLM 没有调用你定义的工具:

原因解决
工具描述不清晰优化 description
参数定义错误检查 JSON Schema
模型不支持使用支持工具的模型
Prompt 不明确明确告诉模型可以使用工具
# 好的工具定义示例
tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        # description 要清晰说明何时使用
        "description": "获取指定城市的当前天气信息。当用户询问天气、温度、是否下雨等问题时使用此工具。",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {
                    "type": "string",
                    # 参数描述也要清晰
                    "description": "城市名称,如:北京、上海、广州"
                },
                "unit": {
                    "type": "string",
                    "enum": ["celsius", "fahrenheit"],
                    "description": "温度单位,默认摄氏度"
                }
            },
            "required": ["city"]
        }
    }
}]

参数解析错误

import json

# 安全解析工具参数
def parse_tool_args(tool_call):
    try:
        return json.loads(tool_call.function.arguments)
    except json.JSONDecodeError as e:
        print(f"参数解析失败: {e}")
        print(f"原始参数: {tool_call.function.arguments}")
        return {}

# 使用
tool_call = response.choices[0].message.tool_calls[0]
args = parse_tool_args(tool_call)

工具调用循环

Agent 陷入无限工具调用循环:

# 限制最大调用次数
MAX_TOOL_CALLS = 10

def run_agent(messages, tools):
    for i in range(MAX_TOOL_CALLS):
        response = client.chat.completions.create(
            model="gpt-4o",
            messages=messages,
            tools=tools
        )
        
        if response.choices[0].finish_reason == "stop":
            return response.choices[0].message.content
        
        # 处理工具调用...
        
    raise Exception("达到最大工具调用次数")

网络问题

代理设置

import httpx
from openai import OpenAI

# 方法1:通过 httpx 设置代理
client = OpenAI(
    http_client=httpx.Client(
        proxies="http://127.0.0.1:7890"
    )
)

# 方法2:环境变量
# export HTTP_PROXY=http://127.0.0.1:7890
# export HTTPS_PROXY=http://127.0.0.1:7890
# export ALL_PROXY=socks5://127.0.0.1:7890
# 测试代理是否工作
curl -x http://127.0.0.1:7890 https://api.openai.com/v1/models \
  -H "Authorization: Bearer $OPENAI_API_KEY"

SSL 证书错误

Error: SSL certificate verify failed
import httpx

# 临时禁用验证(仅用于调试,不推荐生产使用)
client = OpenAI(
    http_client=httpx.Client(verify=False)
)

# 更好的方案:指定证书
client = OpenAI(
    http_client=httpx.Client(verify="/path/to/cert.pem")
)

DNS 解析失败

# 使用自定义 DNS 或直接 IP
import httpx

# 指定 base_url
client = OpenAI(
    base_url="https://api.openai.com/v1",  # 或使用 IP
    http_client=httpx.Client(
        transport=httpx.HTTPTransport(
            retries=3
        )
    )
)

调试技巧

1. 启用日志

import logging

# 设置日志级别
logging.basicConfig(level=logging.DEBUG)

# OpenAI SDK 专用日志
import openai
openai.log = "debug"

2. 打印请求/响应

import httpx

def log_request(request):
    print(f"→ {request.method} {request.url}")
    print(f"  Headers: {dict(request.headers)}")
    if request.content:
        print(f"  Body: {request.content[:500]}...")

def log_response(response):
    print(f"← {response.status_code}")

client = OpenAI(
    http_client=httpx.Client(
        event_hooks={
            "request": [log_request],
            "response": [log_response]
        }
    )
)

3. 使用 curl 测试

# 基础测试
curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

# 带详细输出
curl -v https://api.openai.com/v1/models \
  -H "Authorization: Bearer $OPENAI_API_KEY"

# 测试流式输出
curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Count to 5"}],
    "stream": true
  }'

4. 最小复现

# 创建最小测试用例
from openai import OpenAI

client = OpenAI()

# 最简单的请求
try:
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "Hi"}],
        max_tokens=10
    )
    print("✓ 基础请求成功")
    print(response.choices[0].message.content)
except Exception as e:
    print(f"✗ 错误: {type(e).__name__}: {e}")

错误码速查

错误码含义常见原因解决方案
400Bad Request请求格式错误检查参数格式
401UnauthorizedAPI Key 无效检查 Key
403Forbidden权限不足检查账户权限
404Not Found模型/端点不存在检查模型名称
422Unprocessable参数验证失败检查参数值
429Too Many Requests速率限制降低频率/重试
500Internal Error服务器错误稍后重试
502Bad Gateway网关错误稍后重试
503Service Unavailable服务不可用稍后重试
504Gateway Timeout网关超时增加超时/重试

各厂商状态页:

厂商状态页
OpenAIstatus.openai.com
Anthropicstatus.anthropic.com
Googlestatus.cloud.google.com

相关文档