You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Azure的Web Chat机器人集成Bing Speech API时无法朗读长文本

解决Bing Speech API无法朗读长文本的问题

我之前也碰到过一模一样的问题——Bing Speech API(现在已经整合进Azure Speech Service,但旧的Bing Speech相关逻辑依然适用)对单次文本合成有长度限制,一般单次最多支持1000个字符左右,超过这个长度就会失败或者截断。要解决这个问题,核心思路就是把长文本拆分成符合限制的小片段,然后依次合成播放。

下面是具体的实现方案:

1. 实现长文本拆分函数

我们需要一个函数把长文本拆分成不超过API限制的片段,优先按句子(句号、感叹号、问号)拆分,避免在单词中间截断,保证朗读的连贯性。

// 拆分长文本为符合API限制的片段,默认最大长度设为900(留余量避免触发上限)
function splitLongText(text, maxLength = 900) {
  const chunks = [];
  let currentChunk = '';
  
  // 按标点分割成句子(保留标点符号)
  const sentences = text.split(/([.!?]+)/).filter(Boolean);
  
  for (const sentence of sentences) {
    // 如果当前片段加新句子没超过限制,就合并
    if (currentChunk.length + sentence.length <= maxLength) {
      currentChunk += sentence;
    } else {
      // 把当前片段存入数组
      if (currentChunk) {
        chunks.push(currentChunk.trim());
      }
      // 如果单个句子就超过限制,再按单词拆分
      if (sentence.length > maxLength) {
        const words = sentence.split(' ');
        let tempChunk = '';
        for (const word of words) {
          if (tempChunk.length + word.length + 1 <= maxLength) {
            tempChunk += (tempChunk ? ' ' : '') + word;
          } else {
            chunks.push(tempChunk.trim());
            tempChunk = word;
          }
        }
        if (tempChunk) {
          chunks.push(tempChunk.trim());
        }
      } else {
        currentChunk = sentence;
      }
    }
  }
  
  // 把最后一段存入数组
  if (currentChunk) {
    chunks.push(currentChunk.trim());
  }
  
  return chunks;
}

2. 自定义Web Chat的语音合成逻辑

接下来我们需要替换Web Chat默认的语音合成器,让它处理拆分后的文本片段,逐个播放,确保片段之间连续不重叠。

首先确保你引入了完整的库:

<script src="https://cdn.botframework.com/botframework-webchat/latest/botchat.js"></script>
<script src="https://cdn.botframework.com/botframework-webchat/latest/CognitiveServices.js"></script>

然后配置Web Chat和语音服务:

// 初始化Direct Line连接,替换成你的密钥
const directLine = new BotChat.DirectLine({
  secret: '你的Direct Line密钥'
});

// 自定义语音合成逻辑
const speechOptions = {
  // 保留原有的语音识别配置
  speechRecognizer: new CognitiveServices.SpeechRecognizer({
    subscriptionKey: '你的Bing Speech密钥'
  }),
  // 替换默认的合成器,处理长文本
  speechSynthesizer: {
    synthesize: (text) => {
      return new Promise((resolve, reject) => {
        const textChunks = splitLongText(text);
        let currentChunkIndex = 0;
        
        // 递归播放下一个片段
        function playNextChunk() {
          if (currentChunkIndex >= textChunks.length) {
            resolve();
            return;
          }
          
          const currentChunk = textChunks[currentChunkIndex];
          const synthesizer = new CognitiveServices.SpeechSynthesizer({
            subscriptionKey: '你的Bing Speech密钥'
          });
          
          synthesizer.synthesize(currentChunk)
            .then(() => {
              currentChunkIndex++;
              playNextChunk();
            })
            .catch(error => {
              console.error('合成片段失败:', error);
              reject(error);
            });
        }
        
        playNextChunk();
      });
    }
  }
};

// 渲染Web Chat到页面指定元素
BotChat.App({
  directLine: directLine,
  user: { id: '你的用户ID' },
  bot: { id: '你的机器人ID' },
  speechOptions: speechOptions
}, document.getElementById('webchat'));

额外注意事项

  • 记得把代码里的你的Direct Line密钥、你的Bing Speech密钥、你的用户ID、你的机器人ID替换成你自己的实际值
  • 可以根据实际情况调整maxLength参数,建议不要卡到1000字符的上限,留一些余量避免触发API限制
  • 如果需要更流畅的体验,可以添加一个“正在朗读”的提示,或者处理播放中断的逻辑
  • 另外,Bing Speech API已经被Azure Speech Service取代,建议你后续迁移到新的服务,新服务的文本合成支持更灵活的长度处理,不过上面的拆分逻辑同样适用于新服务

内容的提问来源于stack exchange,提问作者Souvik Ghosh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:01:36