You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于LLM流式输出调用Polly生成语音并在HTML实时播放?

实现流式LLM文本转Polly语音的方案

核心逻辑

通过流式接收LLM输出,实时缓存文本并按「5-6个词」或「标点符号」切割成块,逐块调用AWS Polly生成语音,最后通过播放队列保证语音片段无缝连续播放。


1. 流式接收LLM输出

LLM通常通过Server-Sent Events(SSE)或WebSocket推送流式内容,这里以SSE为例实现前端接收:

// 连接LLM流式接口
const llmStream = new EventSource('/api/llm-stream');
let textBuffer = ''; // 缓存未处理的文本

llmStream.onmessage = (event) => {
  // 累加新收到的文本片段
  textBuffer += event.data;
  // 立即处理缓冲区文本
  processTextBuffer();
};

llmStream.onerror = (error) => {
  console.error('LLM流连接异常:', error);
  llmStream.close();
};

2. 文本块切割逻辑

实现缓冲区文本的切割,优先按标点(中文标点为主)分割,若未达到标点则按5-6个词分割:

function processTextBuffer() {
  // 匹配中文标点符号
  const punctuationRegex = /[。!?,;:]/;
  // 匹配至少5个连续单词(按空格分隔)
  const wordCountRegex = /(\S+\s+){4,5}\S+/;

  let cutIndex = -1;

  // 优先检查是否存在标点
  const punctuationMatch = textBuffer.match(punctuationRegex);
  if (punctuationMatch) {
    cutIndex = textBuffer.indexOf(punctuationMatch[0]) + 1;
  } else {
    // 再检查词数是否达标
    const wordMatch = textBuffer.match(wordCountRegex);
    if (wordMatch) {
      cutIndex = wordMatch[0].length;
    }
  }

  // 提取符合条件的文本块,剩余文本留待后续处理
  if (cutIndex > 0) {
    const textChunk = textBuffer.slice(0, cutIndex);
    textBuffer = textBuffer.slice(cutIndex);
    // 调用Polly生成语音
    generateSpeechChunk(textChunk);
  }
}

3. 调用AWS Polly生成语音流

注意:不要在前端直接暴露AWS密钥,建议通过后端代理Polly接口。以下示例为后端代理后的前端调用逻辑:

// 语音播放队列,保证片段播放顺序
const speechQueue = [];
let isPlaying = false;

async function generateSpeechChunk(text) {
  try {
    // 调用自己的后端代理接口,获取Polly生成的音频Blob
    const response = await fetch('/api/polly-synthesize', {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify({ text, voiceId: 'Zhiyu' })
    });

    if (!response.ok) throw new Error('语音生成失败');
    const audioBlob = await response.blob();
    const audioUrl = URL.createObjectURL(audioBlob);

    // 将音频URL加入播放队列
    speechQueue.push(audioUrl);
    // 若当前无播放任务,启动播放
    if (!isPlaying) playNextChunk();
  } catch (err) {
    console.error('生成语音出错:', err);
  }
}

后端代理Polly的示例(Node.js):

const { PollyClient, SynthesizeSpeechCommand } = require("@aws-sdk/client-polly");
const pollyClient = new PollyClient({ region: 'us-east-1' });

app.post('/api/polly-synthesize', async (req, res) => {
  const { text, voiceId = 'Zhiyu' } = req.body;
  const params = {
    Text: text,
    OutputFormat: 'mp3',
    VoiceId: voiceId
  };

  try {
    const command = new SynthesizeSpeechCommand(params);
    const response = await pollyClient.send(command);
    // 将Polly返回的流直接转发给前端
    response.AudioStream.pipe(res);
  } catch (err) {
    res.status(500).json({ error: err.message });
  }
});

4. 无缝播放语音片段

通过监听音频播放结束事件,依次播放队列中的语音片段,实现连续无卡顿播放:

function playNextChunk() {
  if (speechQueue.length === 0) {
    isPlaying = false;
    return;
  }

  isPlaying = true;
  const audioUrl = speechQueue.shift();
  const audio = new Audio(audioUrl);

  audio.onended = () => {
    URL.revokeObjectURL(audioUrl); // 释放URL资源
    playNextChunk(); // 播放下一个片段
  };

  audio.onerror = (err) => {
    console.error('语音播放失败:', err);
    URL.revokeObjectURL(audioUrl);
    playNextChunk();
  };

  audio.play();
}

关键注意事项

  • 密钥安全:绝对禁止在前端代码中硬编码AWS密钥,必须通过后端代理Polly请求。
  • 切割逻辑调整:可根据LLM输出的实际格式(如是否带空格、标点类型)优化正则表达式。
  • 限流处理:AWS Polly有调用频率限制,若LLM输出过快,可适当增加文本块长度或添加短延迟避免触发限流。

内容的提问来源于stack exchange,提问作者Ouroboros

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 06:00:00