You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取Azure Speech Service TTS输出时长,优化WebChat闲置提示时机?

解决WebChat中语音播报完成后触发闲置提示的问题

针对你遇到的Bot语音播报时误发闲置事件的问题,有两种可靠的方案来实现「语音播报完成n秒后触发闲置提示」的需求,下面分别说明:

方案一:通过Azure Speech SDK获取TTS音频时长

如果你是直接集成Azure Speech SDK处理WebChat的语音合成,可以利用SDK返回的合成结果获取音频时长,再结合活动接收时间计算闲置触发时间:

  1. 自定义语音合成逻辑:在WebChat的配置中重写synthesizeSpeech函数,调用Speech SDK完成文本转语音时,从SynthesisResult中提取audioDuration属性(注意单位是ticks,10000 ticks = 1毫秒)。
  2. 关联时长与消息活动:把每条活动对应的音频时长存储起来(比如用Map),后续在闲置逻辑中使用这个时长计算触发时间——即「活动接收时间 + 音频时长 + n秒」。

示例代码片段:

import { createDirectLine, renderWebChat } from 'botframework-webchat';
import * as SpeechSDK from 'microsoft-cognitiveservices-speech-sdk';

const speechConfig = SpeechSDK.SpeechConfig.fromSubscription('你的订阅密钥', '你的区域');
const synthesizer = new SpeechSDK.SpeechSynthesizer(speechConfig);
const activityAudioDurations = new Map(); // 存储活动ID到音频时长的映射

const webChatOptions = {
  directLine: createDirectLine({ token: '你的DirectLine令牌' }),
  speechOptions: {
    synthesizeSpeech: async (activity) => {
      const result = await synthesizer.speakTextAsync(activity.text);
      
      if (result.reason === SpeechSDK.ResultReason.SynthesizingAudioCompleted) {
        // 转换时长为毫秒
        const audioDurationMs = result.audioDuration / 10000;
        activityAudioDurations.set(activity.id, audioDurationMs);
      }

      return result.audioData;
    }
  }
};

renderWebChat(webChatOptions, document.getElementById('webchat'));

之后在你的闲置事件触发逻辑中,不要直接用活动的接收时间,而是取出对应活动的音频时长,计算出正确的延迟后再触发。

方案二:监听语音播放结束事件(更推荐)

这种方式不需要提前获取时长,而是直接监听语音播放完成的事件,在播放结束后启动n秒倒计时,能更准确地应对网络延迟、播放卡顿等实际场景:

  1. 监听音频播放结束事件:WebChat播放语音时会创建<audio>元素,你可以监听全局的ended事件,判断触发事件的元素是音频元素后,启动倒计时。
  2. 处理多消息播放队列:如果Bot连续发送多条消息,语音会依次播放,这时需要维护一个播放状态标记,确保所有语音播放完成后再启动倒计时。

示例代码片段:

let isPlayingSpeech = false;
let pendingIdleTimeout;

// 监听音频播放结束事件
document.addEventListener('ended', (event) => {
  if (event.target.tagName === 'AUDIO') {
    event.target.removeAttribute('playing');
    // 检查是否还有其他正在播放的音频
    const activeAudios = document.querySelectorAll('audio[playing]');
    isPlayingSpeech = activeAudios.length > 0;

    if (!isPlayingSpeech) {
      // 所有语音播放完成,启动n秒倒计时
      clearTimeout(pendingIdleTimeout);
      pendingIdleTimeout = setTimeout(() => {
        // 发送闲置事件的逻辑
        sendIdleEvent();
      }, n * 1000); // n是你设定的秒数
    }
  }
}, true);

// 在收到新的语音消息开始播放时,标记播放状态
document.addEventListener('play', (event) => {
  if (event.target.tagName === 'AUDIO') {
    event.target.setAttribute('playing', 'true');
    isPlayingSpeech = true;
    // 清除之前的闲置倒计时
    clearTimeout(pendingIdleTimeout);
  }
}, true);

这种方式完全基于实际的播放状态,能避免Bot还在播报时就触发闲置事件的问题,适用性更强。

内容的提问来源于stack exchange,提问作者Manish Sridhar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 00:54:08