Google Speech-to-Text single_utterance模式超时设置问题求助
Google Speech-to-Text single_utterance模式无语音超时解决方案
方法一:使用API原生的speech_start_timeout参数
你之前混淆了参数作用,speech_end_timeout是语音结束后等待后续输入的超时,而**speech_start_timeout才是专门设置等待语音开始的超时时间**。配合enable_voice_activity_events启用语音活动检测,就能在指定时间内未检测到语音时,让API主动结束流式请求,避免程序卡死。
配置示例:
from google.cloud import speech_v1p1beta1 as speech client = speech.SpeechClient() recog_config = speech.RecognitionConfig( encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=16000, language_code="zh-CN", single_utterance=True, ) streaming_config = speech.StreamingRecognitionConfig( config=recog_config, interim_results=False, enable_voice_activity_events=True, speech_start_timeout=30.0 # 30秒内未检测到语音则触发超时 )
超时触发时,API会返回包含语音活动事件的响应或抛出对应RPC错误,你可以捕获该情况结束循环。
方法二:客户端本地定时器兜底
如果API原生参数不生效,可在启动流式识别的同时,启动本地定时器,超时后主动取消流式请求。
代码示例(Python):
import threading from google.cloud import speech_v1p1beta1 as speech def handle_transcription(): client = speech.SpeechClient() recog_config = speech.RecognitionConfig( encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=16000, language_code="zh-CN", single_utterance=True, ) streaming_config = speech.StreamingRecognitionConfig( config=recog_config, interim_results=False ) # 超时回调:主动取消流式请求 def timeout_cancel(stream): print("30秒无语音输入,结束识别") stream.cancel() # 替换为你的实际麦克风音频流生成逻辑 audio_generator = your_microphone_audio_generator() requests = (speech.StreamingRecognizeRequest(audio_content=chunk) for chunk in audio_generator) responses = client.streaming_recognize(streaming_config, requests) # 启动30秒定时器 timer = threading.Timer(30.0, timeout_cancel, args=[responses]) timer.start() try: for response in responses: # 收到识别结果,立即取消定时器 timer.cancel() for result in response.results: print(f"转写结果:{result.alternatives[0].transcript}") except Exception as e: print(f"识别结束:{str(e)}") finally: # 确保定时器被取消,避免内存泄漏 timer.cancel()
方法三:捕获RPC超时异常
流式识别过程中,若API因超时断开连接,会抛出gRPC的DEADLINE_EXCEEDED异常,捕获该异常即可结束循环,避免程序卡死。
示例代码:
from grpc import RpcError, StatusCode try: for response in responses: for result in response.results: print(f"转写结果:{result.alternatives[0].transcript}") except RpcError as e: if e.code() == StatusCode.DEADLINE_EXCEEDED: print("无语音输入,超时退出") else: # 其他异常正常抛出 raise e
内容的提问来源于stack exchange,提问作者Antonio Bono
相关产品推荐
相关产品推荐

