AI流式响应格式化故障求助:修复字符拆分与空格异常
修复AI流式响应文本拆分与格式混乱问题
问题现象
处理AI模型流式响应时,出现以下异常:
- 名称(如Mehran、AIC)、缩写(如You're)被错误拆分
- 标点与单词间存在多余空格,文本显示杂乱
当前处理代码
const { value, done: readerDone } = await reader.read(); done = readerDone; if (value) { let chunk = decoder.decode(value, { stream: true }); chunk = chunk .replace(/0:/g, '') // Remove the '0:' prefix .replace(/"/g, '') // Remove double quotes .replace(/d:\{.*?\}/g, '') // Remove any extra metadata .replace(/\\n\\n/g, '\n\n') // Replace escaped newlines with actual newlines .replace(/}\s*$/, '') // Remove trailing curly bracket .replace(/\s+/g, ' ') // Replace multiple spaces with a single space // Handle markdown-like formatting chunk = chunk .replace(/ \*\s*/g, '\n• ') // Replace bullet points (e.g., "* text") with actual bullet points .replace(/^\s*\*\s+/gm, '• ') // Ensure bullet points at the start of lines .replace(/\n{3,}/g, '\n\n'); // Ensure proper paragraph breaks aiMessageContent += chunk; }
错误格式化示例
"Hello ! Welcome to the A IC Society at Mehr an University of Engineering and Technology ! I 'm excited to chat with you about our events ! We have a wide range of activities and programs organized throughout the year . Can you please tell me what type of events you 're interested in learning more about ? Are you looking for tech -related events , or perhaps something more recreational ?"
原始未格式化响应
"0:"Hello" 0:"!" 0:" Welcome" 0:" to" 0:" the" 0:" A" 0:"IC" 0:" Society" 0:" at" 0:" Mehr" 0:"an" 0:" University" 0:" of" 0:" Engineering" 0:" and" 0:" Technology" 0:"!\n\n" 0:"I" 0:"'m" 0:" excited" 0:" to" 0:" chat" 0:" with" 0:" you" 0:" about" 0:" our" 0:" events" 0:"!" 0:" We" 0:" have" 0:" a" 0:" wide" 0:" range" 0:" of" 0:" activities" 0:" and" 0:" programs" 0:" organized" 0:" throughout" 0:" the" 0:" year" 0:"." 0:" Can" 0:" you" 0:" please" 0:" tell" 0:" me" 0:" what" 0:" type" 0:" of" 0:" events" 0:" you" 0:"'re" 0:" interested" 0:" in" 0:" learning" 0:" more" 0:" about" 0:"?" 0:" Are" 0:" you" 0:" looking" 0:" for" 0:" tech" 0:"-related" 0:" events" 0:"," 0:" or" 0:" perhaps" 0:" something" 0:" more" 0:" recreational" 0:"?" d:{"finishReason":"stop","usage":{"promptTokens":null,"completionTokens":null}}"
解决方案
问题根源是原始响应以0:"内容"为独立单元,直接替换0:会破坏单元间的连接关系。正确做法是先提取每个单元的内容,再拼接:
修改后的代码
const { value, done: readerDone } = await reader.read(); done = readerDone; if (value) { let chunk = decoder.decode(value, { stream: true }); // 提取所有0:"..."中的内容片段,保留原始完整性 const contentMatches = chunk.match(/0:"([^"]*)"/g) || []; chunk = contentMatches.map(match => { // 取出引号内的原始内容 return match.replace(/0:"([^"]*)"/, '$1'); }).join(''); // 处理换行与元数据 chunk = chunk .replace(/\\n\\n/g, '\n\n') // 替换转义换行符为实际换行 .replace(/d:\{.*?\}/g, '') // 移除末尾的元数据块 .replace(/\n{3,}/g, '\n\n'); // 合并连续多余换行 // 处理Markdown项目符号格式 chunk = chunk .replace(/^\s*\*\s+/gm, '• ') // 将行首的*替换为项目符号 .replace(/ \*\s*/g, '\n• '); // 将段落中的*替换为换行项目符号 aiMessageContent += chunk; }
关键修改说明
- 提取完整内容单元:用正则
/0:"([^"]*)"/g精准匹配每个0:"内容"单元,避免直接替换0:导致的片段拆分。 - 正确拼接内容:将提取到的片段直接拼接,自动还原
A+IC=AIC、Mehr+an=Mehran、I+'m=I'm的正确格式。 - 修复标点空格问题:原始响应中标点为独立单元,直接拼接后自动形成
Hello!而非Hello !,解决多余空格。 - 精简无效操作:移除了破坏格式的
replace(/\s+/g, ' '),以及已无必要的replace(/"/g, '')等操作。
修复后效果示例
Hello! Welcome to the AIC Society at Mehran University of Engineering and Technology! I'm excited to chat with you about our events! We have a wide range of activities and programs organized throughout the year. Can you please tell me what type of events you're interested in learning more about? Are you looking for tech-related events, or perhaps something more recreational?
内容的提问来源于stack exchange,提问作者ridzz
相关产品推荐
相关产品推荐

