FastAPI向Next.js14前端SSE流式传输OpenAI响应异常求助
问题:FastAPI SSE流式传输OpenAI响应至Next.js 14前端无法逐词输出
我正在构建AI聊天功能,计划通过FastAPI采用Server-Sent Events(SSE)将OpenAI响应流式传输至Next.js 14(App Router)前端,但实际运行时流要么在首个数据块后停止,要么一次性返回全部数据块,无法像ChatGPT那样实现逐词渐进式输出。
后端(FastAPI)代码
from fastapi import FastAPI from fastapi.responses import StreamingResponse from openai import AsyncOpenAI app = FastAPI() client = AsyncOpenAI(api_key="your-api-key") @app.post("/chat") async def chat(prompt: str): async def generate(): stream = await client.chat.completions.create( model="gpt-4", messages=[{"role": "user", "content": prompt}], stream=True ) async for chunk in stream: delta = chunk.choices[0].delta.content if delta: yield f"data: {delta}\n\n" return StreamingResponse( generate(), media_type="text/event-stream", headers={ "Cache-Control": "no-cache", "X-Accel-Buffering": "no" } )
前端(Next.js 14 App Router)代码
const response = await fetch("/api/chat", { method: "POST", body: JSON.stringify({ prompt }), }); const reader = response.body?.getReader(); const decoder = new TextDecoder(); while (true) { const { done, value } = await reader!.read(); if (done) break; const chunk = decoder.decode(value); console.log(chunk); }
已尝试的解决方案
- 添加
Cache-Control: no-cache请求头 - 添加
X-Accel-Buffering: no请求头 - 尝试结合
fetch()与ReadableStream - 尝试使用EventSource API(但不支持POST请求)
预期与实际效果
- 预期效果:像ChatGPT一样逐词渐进式流式输出内容
- 实际效果:所有文本在完成生成后一次性返回
环境信息
- Python 3.11
- FastAPI 0.110
- openai 1.30
- Next.js 14.2 (App Router)
- 部署环境:Vercel(前端)+ Railway(后端)
解决方案
后端调整
强化SSE头与缓冲区控制
部分部署环境(如Railway)需要更明确的分块传输配置,同时添加事件循环切换逻辑确保数据实时发送:from fastapi import FastAPI from fastapi.responses import StreamingResponse from openai import AsyncOpenAI import asyncio app = FastAPI() client = AsyncOpenAI(api_key="your-api-key") @app.post("/chat") async def chat(prompt: str): async def generate(): stream = await client.chat.completions.create( model="gpt-4", messages=[{"role": "user", "content": prompt}], stream=True ) async for chunk in stream: delta = chunk.choices[0].delta.content if delta: yield f"data: {delta}\n\n" # 触发事件循环切换,强制推送数据 await asyncio.sleep(0) return StreamingResponse( generate(), media_type="text/event-stream", headers={ "Cache-Control": "no-cache, no-store", "X-Accel-Buffering": "no", "Connection": "keep-alive", "Transfer-Encoding": "chunked" } )新增
Connection: keep-alive和Transfer-Encoding: chunked头,明确告知服务器采用分块传输;asyncio.sleep(0)会触发事件循环切换,确保每个生成的内容块被立即发送。验证OpenAI流式响应有效性
在后端生成器中添加打印逻辑,确认OpenAI确实在逐块返回数据:async for chunk in stream: delta = chunk.choices[0].delta.content if delta: print(delta) # 检查控制台是否实时输出单个词/短语 yield f"data: {delta}\n\n" await asyncio.sleep(0)
前端调整
正确解析SSE格式数据流
浏览器接收的数据流可能包含多个SSE消息块,需要按SSE规范拆分解析,避免合并处理:const response = await fetch("/api/chat", { method: "POST", headers: { "Content-Type": "application/json", }, body: JSON.stringify({ prompt }), cache: "no-store", }); if (!response.body) return; const reader = response.body.getReader(); const decoder = new TextDecoder(); let buffer = ""; while (true) { const { done, value } = await reader.read(); if (done) break; // 流式解码,保留未完成的缓冲区内容 buffer += decoder.decode(value, { stream: true }); // 按SSE分隔符拆分消息 const messages = buffer.split(/\n\n/); buffer = messages.pop() || ""; for (const msg of messages) { if (msg.startsWith("data: ")) { const content = msg.slice(6).trim(); console.log(content); // 此处替换为UI更新逻辑,如追加到聊天框 } } }使用
{ stream: true }参数解码,确保不丢失部分数据;按\n\n拆分消息并提取data:前缀后的内容,实现逐词输出。Next.js API路由透传配置
如果前端通过Next.js的/api路由转发请求,需确保路由直接透传流式响应,禁用缓冲:// app/api/chat/route.ts export async function POST(req: Request) { const { prompt } = await req.json(); const response = await fetch(process.env.BACKEND_URL + "/chat", { method: "POST", headers: { "Content-Type": "application/json", }, body: JSON.stringify({ prompt }), cache: "no-store", }); return new Response(response.body, { headers: { "Content-Type": "text/event-stream", "Cache-Control": "no-cache, no-store", "X-Accel-Buffering": "no", "Connection": "keep-alive", "Transfer-Encoding": "chunked", }, }); }
部署环境优化
- Railway后端:移除FastAPI的GZip中间件(避免压缩流式响应),检查应用设置确保无额外缓存策略;
- Vercel前端:确保fetch请求添加
cache: "no-store",禁用Vercel的边缘缓存。
内容的提问来源于stack exchange,提问作者Waqas Ahmed
相关产品推荐
相关产品推荐

