You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用speech_recognition转录2分钟WAV文件时内容截断的解决方法

解决2分钟WAV文件Google语音转录中途截断问题

问题背景

尝试将一段约2分钟的WAV文件转录为文本时,生成的内容中途被截断。已调整识别器的pause_threshold、dynamic_energy_threshold和energy_threshold参数,但问题未解决。

执行环境

  • Python 3.10.13
  • speechrecognition 3.10.0

原测试代码

import speech_recognition as sr

r = sr.Recognizer()
r.pause_threshold = 5
r.dynamic_energy_threshold = False
r.energy_threshold = 200

audio_file = 'horseracing-commentary/data/wav/02_DoncasterMile.wav'

with sr.AudioFile(audio_file) as source:
    audio_data = r.record(source)

# recognize speech using Google Speech Recognition   
try:
    text_data = r.recognize_google(audio_data, language='en-US', show_all=False)

    file_name = audio_file.replace('wav', 'txt')
    f = open(file_name, 'w')
    f.write(text_data)
    f.close()

# Raises a speech_recognition.UnknownValueError exception if the speech is unintelligible.
except sr.UnknownValueError:
    print('Google Speech Recognition could not understand audio')
# Raises a speech_recognition.RequestError exception if the speech recognition operation failed, if the key isn't valid, or if there is no internet connection.
except sr.RequestError as e:
    print('Could not request results from Google Speech Recognition service; {0}'.format(e))

已尝试的无效设置

r.pause_threshold = 5
r.dynamic_energy_threshold = False
r.energy_threshold = 200

解决方案:分块处理长音频

Google免费语音识别API对单次提交的音频时长存在限制(通常约1分钟),直接提交2分钟音频会触发截断。可将音频按固定时长分块,逐块转录后拼接完整结果。

修改后的代码:

import speech_recognition as sr

r = sr.Recognizer()
r.pause_threshold = 5
r.dynamic_energy_threshold = False
r.energy_threshold = 200

audio_file = 'horseracing-commentary/data/wav/02_DoncasterMile.wav'
chunk_duration = 50  # 每块50秒,可根据实际情况调整

full_text = []

with sr.AudioFile(audio_file) as source:
    total_duration = source.DURATION
    current_offset = 0
    
    while current_offset < total_duration:
        # 读取当前时间段的音频数据
        audio_data = r.record(source, duration=chunk_duration, offset=current_offset)
        try:
            # 转录当前音频块
            chunk_text = r.recognize_google(audio_data, language='en-US')
            full_text.append(chunk_text)
        except sr.UnknownValueError:
            print(f"无法识别偏移{current_offset}秒处的音频块")
            full_text.append("[无法识别]")
        except sr.RequestError as e:
            print(f"请求Google语音识别服务失败: {e}")
            break
        # 偏移量递进,处理下一块
        current_offset += chunk_duration

# 拼接所有块的文本并保存
if full_text:
    file_name = audio_file.replace('wav', 'txt')
    with open(file_name, 'w') as f:
        f.write(' '.join(full_text))
    print("转录完成,结果已保存")

额外排查建议

  1. 验证音频完整性:用音频播放器检查WAV文件,确认全程无损坏、无异常静音段。
  2. 查看完整识别响应:将recognize_google的show_all参数设为True,打印返回的完整结果,确认是否有未提取的内容:
result = r.recognize_google(audio_data, language='en-US', show_all=True)
print(result)

如果返回的字典中包含完整转录内容,需调整代码从响应中解析所有结果。


内容的提问来源于stack exchange,提问作者panco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 15:03:16