You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Speech-to-Text single_utterance模式超时设置问题求助

Google Speech-to-Text single_utterance模式无语音超时解决方案

方法一:使用API原生的speech_start_timeout参数

你之前混淆了参数作用,speech_end_timeout是语音结束后等待后续输入的超时,而**speech_start_timeout才是专门设置等待语音开始的超时时间**。配合enable_voice_activity_events启用语音活动检测,就能在指定时间内未检测到语音时,让API主动结束流式请求,避免程序卡死。

配置示例:

from google.cloud import speech_v1p1beta1 as speech

client = speech.SpeechClient()
recog_config = speech.RecognitionConfig(
    encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16,
    sample_rate_hertz=16000,
    language_code="zh-CN",
    single_utterance=True,
)

streaming_config = speech.StreamingRecognitionConfig(
    config=recog_config,
    interim_results=False,
    enable_voice_activity_events=True,
    speech_start_timeout=30.0  # 30秒内未检测到语音则触发超时
)

超时触发时,API会返回包含语音活动事件的响应或抛出对应RPC错误,你可以捕获该情况结束循环。

方法二:客户端本地定时器兜底

如果API原生参数不生效,可在启动流式识别的同时,启动本地定时器,超时后主动取消流式请求。

代码示例(Python):

import threading
from google.cloud import speech_v1p1beta1 as speech

def handle_transcription():
    client = speech.SpeechClient()
    recog_config = speech.RecognitionConfig(
        encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16,
        sample_rate_hertz=16000,
        language_code="zh-CN",
        single_utterance=True,
    )
    streaming_config = speech.StreamingRecognitionConfig(
        config=recog_config,
        interim_results=False
    )

    # 超时回调:主动取消流式请求
    def timeout_cancel(stream):
        print("30秒无语音输入,结束识别")
        stream.cancel()

    # 替换为你的实际麦克风音频流生成逻辑
    audio_generator = your_microphone_audio_generator()
    requests = (speech.StreamingRecognizeRequest(audio_content=chunk) for chunk in audio_generator)

    responses = client.streaming_recognize(streaming_config, requests)
    # 启动30秒定时器
    timer = threading.Timer(30.0, timeout_cancel, args=[responses])
    timer.start()

    try:
        for response in responses:
            # 收到识别结果,立即取消定时器
            timer.cancel()
            for result in response.results:
                print(f"转写结果:{result.alternatives[0].transcript}")
    except Exception as e:
        print(f"识别结束:{str(e)}")
    finally:
        # 确保定时器被取消,避免内存泄漏
        timer.cancel()

方法三:捕获RPC超时异常

流式识别过程中,若API因超时断开连接,会抛出gRPC的DEADLINE_EXCEEDED异常,捕获该异常即可结束循环,避免程序卡死。

示例代码:

from grpc import RpcError, StatusCode

try:
    for response in responses:
        for result in response.results:
            print(f"转写结果:{result.alternatives[0].transcript}")
except RpcError as e:
    if e.code() == StatusCode.DEADLINE_EXCEEDED:
        print("无语音输入,超时退出")
    else:
        # 其他异常正常抛出
        raise e

内容的提问来源于stack exchange,提问作者Antonio Bono

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 05:12:29