使用speech_recognition转录2分钟WAV文件时内容截断的解决方法
解决2分钟WAV文件Google语音转录中途截断问题
问题背景
尝试将一段约2分钟的WAV文件转录为文本时,生成的内容中途被截断。已调整识别器的pause_threshold、dynamic_energy_threshold和energy_threshold参数,但问题未解决。
执行环境
- Python 3.10.13
- speechrecognition 3.10.0
原测试代码
import speech_recognition as sr r = sr.Recognizer() r.pause_threshold = 5 r.dynamic_energy_threshold = False r.energy_threshold = 200 audio_file = 'horseracing-commentary/data/wav/02_DoncasterMile.wav' with sr.AudioFile(audio_file) as source: audio_data = r.record(source) # recognize speech using Google Speech Recognition try: text_data = r.recognize_google(audio_data, language='en-US', show_all=False) file_name = audio_file.replace('wav', 'txt') f = open(file_name, 'w') f.write(text_data) f.close() # Raises a speech_recognition.UnknownValueError exception if the speech is unintelligible. except sr.UnknownValueError: print('Google Speech Recognition could not understand audio') # Raises a speech_recognition.RequestError exception if the speech recognition operation failed, if the key isn't valid, or if there is no internet connection. except sr.RequestError as e: print('Could not request results from Google Speech Recognition service; {0}'.format(e))
已尝试的无效设置
r.pause_threshold = 5 r.dynamic_energy_threshold = False r.energy_threshold = 200
解决方案:分块处理长音频
Google免费语音识别API对单次提交的音频时长存在限制(通常约1分钟),直接提交2分钟音频会触发截断。可将音频按固定时长分块,逐块转录后拼接完整结果。
修改后的代码:
import speech_recognition as sr r = sr.Recognizer() r.pause_threshold = 5 r.dynamic_energy_threshold = False r.energy_threshold = 200 audio_file = 'horseracing-commentary/data/wav/02_DoncasterMile.wav' chunk_duration = 50 # 每块50秒,可根据实际情况调整 full_text = [] with sr.AudioFile(audio_file) as source: total_duration = source.DURATION current_offset = 0 while current_offset < total_duration: # 读取当前时间段的音频数据 audio_data = r.record(source, duration=chunk_duration, offset=current_offset) try: # 转录当前音频块 chunk_text = r.recognize_google(audio_data, language='en-US') full_text.append(chunk_text) except sr.UnknownValueError: print(f"无法识别偏移{current_offset}秒处的音频块") full_text.append("[无法识别]") except sr.RequestError as e: print(f"请求Google语音识别服务失败: {e}") break # 偏移量递进,处理下一块 current_offset += chunk_duration # 拼接所有块的文本并保存 if full_text: file_name = audio_file.replace('wav', 'txt') with open(file_name, 'w') as f: f.write(' '.join(full_text)) print("转录完成,结果已保存")
额外排查建议
- 验证音频完整性:用音频播放器检查WAV文件,确认全程无损坏、无异常静音段。
- 查看完整识别响应:将
recognize_google的show_all参数设为True,打印返回的完整结果,确认是否有未提取的内容:
result = r.recognize_google(audio_data, language='en-US', show_all=True) print(result)
如果返回的字典中包含完整转录内容,需调整代码从响应中解析所有结果。
内容的提问来源于stack exchange,提问作者panco
相关产品推荐
相关产品推荐

