You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用Whisper模块音频转文本耗时过长的优化问询

缩短Whisper音频转文本处理时间的优化方案

核心优化措施

1. 避免重复加载模型

你当前代码每次调用transcribe_audio都会重新加载tiny模型,这是主要的耗时点之一。把模型加载移到函数外部,全局初始化一次即可:

# 全局加载模型,仅执行一次
model = whisper.load_model('tiny')

def transcribe_audio(filename): 
    print(filename)
    location = r"C:\Users\shree\Desktop\projects\Summarization_project\Point_Extraction\media\audio_files"
    file_path = os.path.join(location, filename)
    result = model.transcribe(file_path, fp16=False)
    text = result['text']
    text = auto_correct_text(text)
    print(text)
    os.remove(file_path)
    print(f"Deleted: {filename}")
    return text

2. 启用硬件加速(GPU)

如果你的设备有NVIDIA GPU,开启半精度推理(fp16=True),能大幅提升处理速度,默认该参数为True,你当前手动设为False了:

# 确保已安装GPU版本PyTorch,修改transcribe参数
result = model.transcribe(file_path, fp16=True)

无GPU设备可忽略此优化,保持fp16=False即可

3. 指定音频语言

明确设置language参数(如英文"en"、中文"zh"),跳过模型自动检测语言的步骤,节省时间:

result = model.transcribe(file_path, fp16=False, language="en")

4. 音频预处理优化

  • 提前将音频转成Whisper原生支持的格式(如WAV、MP3),避免模型处理时额外做格式转换
  • 裁剪音频中的静音片段,只转写有效音频部分

5. 批量处理音频

如果有多个音频文件,改为批量加载处理,减少文件IO操作的重复开销

内容的提问来源于stack exchange,提问作者vaibhav girase

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 06:32:52