You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Twilio+Google流式STT的电话应用is_final始终为False求助

问题:Google流式STT的is_final始终为False,无法触发完整语句识别

我用Twilio Media Streams搭建电话应用,流程为:Twilio Media Stream → Google STT(流式)→ LLM → TTS。基于官方Python示例代码修改了on_transcription_response函数:

def on_transcription_response(response):
    if not response.results:
        return

    result = response.results[0]
    if not result.alternatives:
        return

    transcription = result.alternatives[0].transcript
    print("Transcription: " + transcription + " is_final: " + str(result.is_final))

遇到的问题是result.is_final从未返回True,导致无法将完整转录内容传入LLM。我尝试添加静音检测过滤静音片段:

import audioop

def is_silence(buffer, threshold=500):
    pcm = audioop.ulaw2lin(buffer, 2)  # 转成16位PCM
    rms = audioop.rms(pcm, 2)          # 计算均方根振幅
    return rms < threshold

def add_request(self, buffer):
    if is_silence(buffer):
        print("Skipping silence based on amplitude")
        return
    self._queue.put(bytes(buffer), block=False)

但is_final仍始终为False。另外我需要用language_code="yue-Hant-HK"持续识别语音,通话过程中要随时检测完整语句,不能识别一次就停止。请问如何解决?


解决方案
  • 调整Google STT流式识别核心参数
    Google STT默认不会主动返回is_final=True,需显式配置参数引导它判断语句边界。初始化流式识别请求时,重点配置以下内容:

    from google.cloud.speech_v2 import SpeechClient
    from google.cloud.speech_v2.types import RecognitionConfig, StreamingRecognitionConfig
    
    recognition_config = RecognitionConfig(
        auto_decoding_config={
            "encoding": "MULAW",
            "sample_rate_hertz": 8000,
            "audio_channel_count": 1
        },
        language_codes=["yue-Hant-HK"],
        model="long",  # 适配长语音通话场景
        features=RecognitionConfig.Features(
            enable_automatic_punctuation=True,  # 自动标点帮助判断语句结束
            enable_word_time_offsets=True
        ),
    )
    streaming_config = StreamingRecognitionConfig(
        config=recognition_config,
        interim_results=True,
        single_utterance=False,  # 关键:设为False实现持续识别,不自动停止
    )
    

    其中single_utterance=False确保通话过程中识别不中断,enable_automatic_punctuation=True让STT通过标点符号判断语句结束,从而返回is_final=True。

  • 修复静音检测逻辑,改为触发结束信号
    直接丢弃静音片段会导致STT无法判断说话停顿,进而不返回is_final=True。正确做法是保留静音,统计连续静音时长,达到阈值时手动向STT发送结束请求:

    from google.cloud.speech_v2.types import StreamingRecognizeRequest
    
    silence_counter = 0
    SILENCE_THRESHOLD_FRAMES = 20  # 按8kHz采样率计算,约1.25秒连续静音
    
    def add_request(self, buffer):
        global silence_counter
        if is_silence(buffer):
            silence_counter += 1
            if silence_counter >= SILENCE_THRESHOLD_FRAMES:
                # 连续静音达标,发送结束信号触发is_final
                self._queue.put(StreamingRecognizeRequest(end_of_single_utterance=True), block=False)
                silence_counter = 0
        else:
            silence_counter = 0
            self._queue.put(bytes(buffer), block=False)
    

    发送结束信号后,因为single_utterance=False,STT会返回当前语句的is_final=True结果,同时继续监听后续音频。

  • 确保音频格式完全匹配
    Twilio Media Streams默认输出8kHz采样率、μLaw编码的单声道音频,必须在Google STT配置中明确指定解码参数(如上面代码中的auto_decoding_config),否则格式不匹配会导致STT无法正确解析音频,无法判断语句边界。

  • 针对粤语识别的优化
    粤语的语音停顿逻辑与普通话不同,可添加粤语常用短语到speech_contexts,帮助STT更精准识别语句边界:

    features=RecognitionConfig.Features(
        # 其他参数...
        speech_contexts=[{"phrases": ["係", "唔係", "多謝", "唔該", "嘅"]}]
    )
    

内容的提问来源于stack exchange,提问作者J7er

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 14:42:03