基于Azure的Web Chat机器人集成Bing Speech API时无法朗读长文本
解决Bing Speech API无法朗读长文本的问题
我之前也碰到过一模一样的问题——Bing Speech API(现在已经整合进Azure Speech Service,但旧的Bing Speech相关逻辑依然适用)对单次文本合成有长度限制,一般单次最多支持1000个字符左右,超过这个长度就会失败或者截断。要解决这个问题,核心思路就是把长文本拆分成符合限制的小片段,然后依次合成播放。
下面是具体的实现方案:
1. 实现长文本拆分函数
我们需要一个函数把长文本拆分成不超过API限制的片段,优先按句子(句号、感叹号、问号)拆分,避免在单词中间截断,保证朗读的连贯性。
// 拆分长文本为符合API限制的片段,默认最大长度设为900(留余量避免触发上限) function splitLongText(text, maxLength = 900) { const chunks = []; let currentChunk = ''; // 按标点分割成句子(保留标点符号) const sentences = text.split(/([.!?]+)/).filter(Boolean); for (const sentence of sentences) { // 如果当前片段加新句子没超过限制,就合并 if (currentChunk.length + sentence.length <= maxLength) { currentChunk += sentence; } else { // 把当前片段存入数组 if (currentChunk) { chunks.push(currentChunk.trim()); } // 如果单个句子就超过限制,再按单词拆分 if (sentence.length > maxLength) { const words = sentence.split(' '); let tempChunk = ''; for (const word of words) { if (tempChunk.length + word.length + 1 <= maxLength) { tempChunk += (tempChunk ? ' ' : '') + word; } else { chunks.push(tempChunk.trim()); tempChunk = word; } } if (tempChunk) { chunks.push(tempChunk.trim()); } } else { currentChunk = sentence; } } } // 把最后一段存入数组 if (currentChunk) { chunks.push(currentChunk.trim()); } return chunks; }
2. 自定义Web Chat的语音合成逻辑
接下来我们需要替换Web Chat默认的语音合成器,让它处理拆分后的文本片段,逐个播放,确保片段之间连续不重叠。
首先确保你引入了完整的库:
<script src="https://cdn.botframework.com/botframework-webchat/latest/botchat.js"></script> <script src="https://cdn.botframework.com/botframework-webchat/latest/CognitiveServices.js"></script>
然后配置Web Chat和语音服务:
// 初始化Direct Line连接,替换成你的密钥 const directLine = new BotChat.DirectLine({ secret: '你的Direct Line密钥' }); // 自定义语音合成逻辑 const speechOptions = { // 保留原有的语音识别配置 speechRecognizer: new CognitiveServices.SpeechRecognizer({ subscriptionKey: '你的Bing Speech密钥' }), // 替换默认的合成器,处理长文本 speechSynthesizer: { synthesize: (text) => { return new Promise((resolve, reject) => { const textChunks = splitLongText(text); let currentChunkIndex = 0; // 递归播放下一个片段 function playNextChunk() { if (currentChunkIndex >= textChunks.length) { resolve(); return; } const currentChunk = textChunks[currentChunkIndex]; const synthesizer = new CognitiveServices.SpeechSynthesizer({ subscriptionKey: '你的Bing Speech密钥' }); synthesizer.synthesize(currentChunk) .then(() => { currentChunkIndex++; playNextChunk(); }) .catch(error => { console.error('合成片段失败:', error); reject(error); }); } playNextChunk(); }); } } }; // 渲染Web Chat到页面指定元素 BotChat.App({ directLine: directLine, user: { id: '你的用户ID' }, bot: { id: '你的机器人ID' }, speechOptions: speechOptions }, document.getElementById('webchat'));
额外注意事项
- 记得把代码里的
你的Direct Line密钥、你的Bing Speech密钥、你的用户ID、你的机器人ID替换成你自己的实际值 - 可以根据实际情况调整
maxLength参数,建议不要卡到1000字符的上限,留一些余量避免触发API限制 - 如果需要更流畅的体验,可以添加一个“正在朗读”的提示,或者处理播放中断的逻辑
- 另外,Bing Speech API已经被Azure Speech Service取代,建议你后续迁移到新的服务,新服务的文本合成支持更灵活的长度处理,不过上面的拆分逻辑同样适用于新服务
内容的提问来源于stack exchange,提问作者Souvik Ghosh
相关产品推荐
相关产品推荐

