如何基于LLM流式输出调用Polly生成语音并在HTML实时播放?
实现流式LLM文本转Polly语音的方案
核心逻辑
通过流式接收LLM输出,实时缓存文本并按「5-6个词」或「标点符号」切割成块,逐块调用AWS Polly生成语音,最后通过播放队列保证语音片段无缝连续播放。
1. 流式接收LLM输出
LLM通常通过Server-Sent Events(SSE)或WebSocket推送流式内容,这里以SSE为例实现前端接收:
// 连接LLM流式接口 const llmStream = new EventSource('/api/llm-stream'); let textBuffer = ''; // 缓存未处理的文本 llmStream.onmessage = (event) => { // 累加新收到的文本片段 textBuffer += event.data; // 立即处理缓冲区文本 processTextBuffer(); }; llmStream.onerror = (error) => { console.error('LLM流连接异常:', error); llmStream.close(); };
2. 文本块切割逻辑
实现缓冲区文本的切割,优先按标点(中文标点为主)分割,若未达到标点则按5-6个词分割:
function processTextBuffer() { // 匹配中文标点符号 const punctuationRegex = /[。!?,;:]/; // 匹配至少5个连续单词(按空格分隔) const wordCountRegex = /(\S+\s+){4,5}\S+/; let cutIndex = -1; // 优先检查是否存在标点 const punctuationMatch = textBuffer.match(punctuationRegex); if (punctuationMatch) { cutIndex = textBuffer.indexOf(punctuationMatch[0]) + 1; } else { // 再检查词数是否达标 const wordMatch = textBuffer.match(wordCountRegex); if (wordMatch) { cutIndex = wordMatch[0].length; } } // 提取符合条件的文本块,剩余文本留待后续处理 if (cutIndex > 0) { const textChunk = textBuffer.slice(0, cutIndex); textBuffer = textBuffer.slice(cutIndex); // 调用Polly生成语音 generateSpeechChunk(textChunk); } }
3. 调用AWS Polly生成语音流
注意:不要在前端直接暴露AWS密钥,建议通过后端代理Polly接口。以下示例为后端代理后的前端调用逻辑:
// 语音播放队列,保证片段播放顺序 const speechQueue = []; let isPlaying = false; async function generateSpeechChunk(text) { try { // 调用自己的后端代理接口,获取Polly生成的音频Blob const response = await fetch('/api/polly-synthesize', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ text, voiceId: 'Zhiyu' }) }); if (!response.ok) throw new Error('语音生成失败'); const audioBlob = await response.blob(); const audioUrl = URL.createObjectURL(audioBlob); // 将音频URL加入播放队列 speechQueue.push(audioUrl); // 若当前无播放任务,启动播放 if (!isPlaying) playNextChunk(); } catch (err) { console.error('生成语音出错:', err); } }
后端代理Polly的示例(Node.js):
const { PollyClient, SynthesizeSpeechCommand } = require("@aws-sdk/client-polly"); const pollyClient = new PollyClient({ region: 'us-east-1' }); app.post('/api/polly-synthesize', async (req, res) => { const { text, voiceId = 'Zhiyu' } = req.body; const params = { Text: text, OutputFormat: 'mp3', VoiceId: voiceId }; try { const command = new SynthesizeSpeechCommand(params); const response = await pollyClient.send(command); // 将Polly返回的流直接转发给前端 response.AudioStream.pipe(res); } catch (err) { res.status(500).json({ error: err.message }); } });
4. 无缝播放语音片段
通过监听音频播放结束事件,依次播放队列中的语音片段,实现连续无卡顿播放:
function playNextChunk() { if (speechQueue.length === 0) { isPlaying = false; return; } isPlaying = true; const audioUrl = speechQueue.shift(); const audio = new Audio(audioUrl); audio.onended = () => { URL.revokeObjectURL(audioUrl); // 释放URL资源 playNextChunk(); // 播放下一个片段 }; audio.onerror = (err) => { console.error('语音播放失败:', err); URL.revokeObjectURL(audioUrl); playNextChunk(); }; audio.play(); }
关键注意事项
- 密钥安全:绝对禁止在前端代码中硬编码AWS密钥,必须通过后端代理Polly请求。
- 切割逻辑调整:可根据LLM输出的实际格式(如是否带空格、标点类型)优化正则表达式。
- 限流处理:AWS Polly有调用频率限制,若LLM输出过快,可适当增加文本块长度或添加短延迟避免触发限流。
内容的提问来源于stack exchange,提问作者Ouroboros
相关产品推荐
相关产品推荐

