流式输出实现详解
SSE 解析、前端展示和打字机效果的完整实现方案。
为什么需要流式输出?
大模型生成响应需要数秒甚至更长时间。流式输出让用户实时看到生成内容,显著提升体验。
| 模式 | 首字延迟 | 用户体验 | 适用场景 |
|---|---|---|---|
| 非流式 | 5-30秒 | 长时间等待,体验差 | 批量处理、后台任务 |
| 流式 | 0.5-2秒 | 实时看到内容,体验好 | 聊天、实时交互 |
流式输出的核心价值:
- 降低感知延迟:用户立即看到响应开始,而非等待完整结果
- 提升交互感:打字机效果让 AI 更像”在思考”
- 支持长响应:避免长时间无响应导致的超时
- 可中断:用户可以随时停止生成
技术原理:
流式输出基于 Server-Sent Events (SSE) 协议。服务器保持连接打开,逐步发送数据块,客户端实时接收并渲染。
SSE 协议回顾
SSE(Server-Sent Events)是 HTML5 标准的一部分,专门用于服务器向客户端推送数据。
sequenceDiagram
participant Client
participant Server
participant LLM
Client->>Server: POST /chat (stream=true)
Server->>LLM: 请求
loop 逐块返回
LLM-->>Server: token
Server-->>Client: data: {...}
end
Server-->>Client: data: [DONE]SSE vs WebSocket:
| 特性 | SSE | WebSocket |
|---|---|---|
| 方向 | 单向(服务器→客户端) | 双向 |
| 协议 | HTTP | 独立协议 |
| 复杂度 | 简单 | 较复杂 |
| 重连 | 自动 | 需手动 |
| 适用场景 | 流式输出 | 实时双向通信 |
对于 LLM 流式输出,SSE 是最佳选择,因为只需要服务器向客户端推送。
数据格式
SSE 使用简单的文本格式,每条消息以 data: 开头,以两个换行符结束:
data: {"choices":[{"delta":{"content":"你"}}]}
data: {"choices":[{"delta":{"content":"好"}}]}
data: [DONE]
格式规范:
- 每行以
data:开头 - 消息之间用空行分隔(
\n\n) - 可选的
event:字段指定事件类型 - 可选的
id:字段用于断点续传
后端实现
Python FastAPI
FastAPI 的 StreamingResponse 非常适合实现 SSE:
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from openai import OpenAI
import json
app = FastAPI()
client = OpenAI()
@app.post("/chat")
async def chat(message: str):
async def generate():
stream = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": message}],
stream=True
)
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
yield f"data: {json.dumps({'content': content})}\n\n"
yield "data: [DONE]\n\n"
return StreamingResponse(
generate(),
media_type="text/event-stream",
headers={
"Cache-Control": "no-cache",
"Connection": "keep-alive",
}
)
关键点:
media_type="text/event-stream"告诉浏览器这是 SSE- 每个数据块以
data:开头,以\n\n结尾 - 使用生成器函数实现流式返回
Node.js Express
const express = require('express');
const OpenAI = require('openai');
const app = express();
const openai = new OpenAI();
app.post('/chat', async (req, res) => {
// 设置 SSE 响应头
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache');
res.setHeader('Connection', 'keep-alive');
const stream = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: req.body.message }],
stream: true
});
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content;
if (content) {
res.write(`data: ${JSON.stringify({ content })}\n\n`);
}
}
res.write('data: [DONE]\n\n');
res.end();
});
错误处理:
try {
for await (const chunk of stream) {
// ...
}
} catch (error) {
res.write(`data: ${JSON.stringify({ error: error.message })}\n\n`);
} finally {
res.end();
}
前端实现
原生 JavaScript
使用 Fetch API 的 ReadableStream 处理流式响应:
async function streamChat(message) {
const response = await fetch('/chat', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ message })
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let result = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
const chunk = decoder.decode(value);
const lines = chunk.split('\n');
for (const line of lines) {
if (line.startsWith('data: ')) {
const data = line.slice(6);
if (data === '[DONE]') return result;
try {
const { content } = JSON.parse(data);
result += content;
updateUI(result); // 更新界面
} catch (e) {
// 忽略解析错误(可能是不完整的 JSON)
}
}
}
}
return result;
}
注意事项:
- 数据可能跨多个 chunk 到达,需要正确处理边界
- JSON 可能被截断,需要错误处理
- 考虑使用专门的 SSE 解析库
React Hook
function useStreamChat() {
const [content, setContent] = useState('');
const [loading, setLoading] = useState(false);
const sendMessage = async (message) => {
setLoading(true);
setContent('');
const response = await fetch('/chat', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ message })
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const chunk = decoder.decode(value);
for (const line of chunk.split('\n')) {
if (line.startsWith('data: ') && line.slice(6) !== '[DONE]') {
const { content: c } = JSON.parse(line.slice(6));
setContent(prev => prev + c);
}
}
}
setLoading(false);
};
return { content, loading, sendMessage };
}
Vue 3 Composable
export function useStreamChat() {
const content = ref('');
const loading = ref(false);
async function sendMessage(message) {
loading.value = true;
content.value = '';
const response = await fetch('/chat', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ message })
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const chunk = decoder.decode(value);
for (const line of chunk.split('\n')) {
if (line.startsWith('data: ') && line.slice(6) !== '[DONE]') {
const { content: c } = JSON.parse(line.slice(6));
content.value += c;
}
}
}
loading.value = false;
}
return { content, loading, sendMessage };
}
打字机效果
CSS 光标闪烁
.typing-cursor::after {
content: '▋';
animation: blink 1s infinite;
}
@keyframes blink {
0%, 50% { opacity: 1; }
51%, 100% { opacity: 0; }
}
平滑滚动
function scrollToBottom(container) {
container.scrollTo({
top: container.scrollHeight,
behavior: 'smooth'
});
}
// 每次内容更新时调用
useEffect(() => {
scrollToBottom(containerRef.current);
}, [content]);
Markdown 渲染
流式输出的 Markdown 需要实时渲染,推荐使用 react-markdown。
import ReactMarkdown from 'react-markdown';
import { Prism as SyntaxHighlighter } from 'react-syntax-highlighter';
function ChatMessage({ content }) {
return (
<ReactMarkdown
components={{
code({ node, inline, className, children, ...props }) {
const match = /language-(\w+)/.exec(className || '');
return !inline && match ? (
<SyntaxHighlighter language={match[1]} {...props}>
{String(children).replace(/\n$/, '')}
</SyntaxHighlighter>
) : (
<code className={className} {...props}>{children}</code>
);
}
}}
>
{content}
</ReactMarkdown>
);
}
错误处理
连接中断恢复
async function streamWithRetry(message, maxRetries = 3) {
let retries = 0;
let accumulated = '';
while (retries < maxRetries) {
try {
const response = await fetch('/chat', {
method: 'POST',
body: JSON.stringify({ message, resume_from: accumulated })
});
// ... 处理流
return accumulated;
} catch (error) {
retries++;
if (retries >= maxRetries) throw error;
await sleep(1000 * retries);
}
}
}
超时处理
const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 60000);
try {
const response = await fetch('/chat', {
signal: controller.signal,
// ...
});
// ...
} finally {
clearTimeout(timeout);
}
性能优化
防抖渲染
// 避免每个 token 都触发渲染
const debouncedUpdate = useMemo(
() => debounce((text) => setDisplayContent(text), 16),
[]
);
// 在流式接收时
content += newToken;
debouncedUpdate(content);
虚拟滚动
长对话使用虚拟滚动避免 DOM 过多:
import { FixedSizeList } from 'react-window';
<FixedSizeList
height={600}
itemCount={messages.length}
itemSize={100}
>
{({ index, style }) => (
<div style={style}>
<ChatMessage message={messages[index]} />
</div>
)}
</FixedSizeList>
完整示例架构
flowchart TD
A[用户输入] --> B[前端]
B --> C[API 网关]
C --> D[后端服务]
D --> E[LLM API]
E -->|SSE| D
D -->|SSE| C
C -->|SSE| B
B --> F[Markdown 渲染]
F --> G[打字机效果]
G --> H[显示给用户]