Azure AI Hub Phi3 mini流式响应数据缺失排查与修复咨询
Azure AI Hub流式获取Phi3 mini响应时内容缺失问题排查与修复
问题背景
在Azure AI Hub无服务器部署环境中,流式调用Phi3 mini模型的聊天补全API时,出现随机字符/单词缺失的情况。关闭流式功能后API运行正常,尝试使用第三方SSE解析器后问题依然存在。
错误输出示例
data: {"id":"cmpl-6efec8012ebb48768a85a7489211369f","object":" AP PHYSICS 2 (3 Sources for Physics Lesson Key Topics to Study: Introduction todata: {"id":"cmpl-6efec8012ebb48768a85a7489211369f"," ources: Descriptions of point, line, and surface sources, field patterns, and direction of fields Radio Waves: Generation, propagation, and applications of radio waves data: {"id":"cmpl-6efec8012ebb48768a85a7489211369f","object":"chat.co and Wave Propagation: Transmission modes, reflection, refraction, and absorption in various media Study Tips: Understanding Source Types:data: {"id":"cmpl-6efec8012ebb48768a85a7489 iarize yourself with diagrams of different source types and their field patterns. Practice Problems: Work through problems related to field strength, intensity, and coverage areas ofdata: {"id":"cmpl-6efec8012e sources. Visual Aids: Use animations and simulations to visualize wave propagation and interactions with various media. Resources: Textbook: Review the physics ofdata: {"id":"cmpl-6efec8012ebb48768a85a7489 and sources chapter in your AP Physics textbook. OpenStax: Access "AP Physics 2: Physics for College Students" by OpenStax for free, whichdata: {"id":"cmpl-6efec8012ebb48768a85a7489211369f","obje a range of problems and explanations. MIT OpenCourseWare: Explore the "AP Physics 2: Physics for College Students" course materials on Mdata: {"id":"cmpl-6efec8012ebb487 's open courseware website, which includes video lectures and problem sets. Remember to align your study efforts with the particular questions and concepts that will likely appear on AP Physics 2 exam, with an emphasis on understanding the principles of wave sources, as these are fundamental to many AP Physics topics.
相关代码
请求构造
const headers = { "Content-Type": "application/json", "Authorization": "Bearer " + env.AI_KEY } const body = { "messages": [ { "role": "user", "content": "..." }, { "role": "user", "content": course + " " + name } ], "max_tokens": 1024, "temperature": 0.8, "top_p": 1, "stream": true } const response = await fetch(endpoint, { method: "POST", headers: headers, body: JSON.stringify(body) })
响应读取器
const SSEEvents = { onError: (error: any) => { console.error(error); controller.error(error) }, onData: (data: string) => { const queue = new TextEncoder().encode(data); process.stdout.write(queue) controller.enqueue(queue); }, onComplete: () => { //console.log("Done") controller.close() }, }; const decoder = new TextDecoder(); const sseParser = new SSEParser(SSEEvents); const reader = response.body?.getReader() while (true) { const book = await reader?.read(); if (book?.done) break; const chunk = decoder.decode(book?.value); try { sseParser.parseSSE(chunk); } catch (error) { console.log(chunk) //console.log(error) } } controller.close()
问题定位
从输出的截断内容(比如"object":"chat.co")来看,问题核心在流式响应的客户端处理逻辑,而非Azure服务端:非流式请求能返回完整内容,说明模型和服务端本身可以生成正确结果,只是流式传输的chunk解码或解析环节出现了截断。
修复方案
1. 修正TextDecoder的流式解码逻辑
当前代码中decoder.decode(book?.value)默认会一次性消费所有字节并重置解码器状态,若chunk包含UTF-8多字节字符的一部分,会导致字符截断。需要添加{ stream: true }参数,让解码器缓存未完成的字节,等待下一个chunk到来后再完整解码:
const chunk = decoder.decode(book?.value, { stream: true });
2. 实现支持增量解析的SSE处理逻辑
第三方SSE解析器可能未考虑事件被拆分到多个chunk的场景,需要自己维护缓冲区处理不完整事件:
- 用全局变量存储未处理的剩余内容
- 每次新chunk到来后,与缓冲区拼接,按SSE事件分隔符
\n\n分割出完整事件 - 处理每个完整事件,剩余的不完整内容留到下次循环处理
示例简化代码:
let buffer = ''; const decoder = new TextDecoder(); const reader = response.body?.getReader(); while (true) { const { done, value } = await reader.read(); if (done) break; buffer += decoder.decode(value, { stream: true }); const events = buffer.split('\n\n'); buffer = events.pop(); // 保留未完成的部分 for (const event of events) { if (!event.startsWith('data: ')) continue; const data = event.slice(6).trim(); if (data === '[DONE]') continue; try { const json = JSON.parse(data); const content = json.choices[0].delta.content; if (content) { // 处理获取到的内容片段 process.stdout.write(content); controller.enqueue(content); } } catch (e) { console.error('解析SSE数据失败:', e); } } } // 处理最后剩余的缓冲区内容(如果有的话) if (buffer) { // 逻辑同上,尝试解析剩余内容 } controller.close();
3. 验证fetch请求配置
确保fetch请求未设置错误的responseType,默认配置即可保证response.body为可读取的ReadableStream,无需额外修改。
参考方向
- 学习SSE规范中事件分块传输的规则,理解事件边界的判断方式
- 熟悉Node.js/浏览器中ReadableStream的读取与字节解码机制
- 参考Azure官方提供的流式响应处理示例,对齐标准处理流程
内容的提问来源于stack exchange,提问作者Techno Talks
相关产品推荐
相关产品推荐

