You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用OpenAI Whisper转录全音频时遇AssertionError错误咨询

解决Whisper转录完整音频时的AssertionError: incorrect audio shape问题

你遇到的问题根源是:Whisper的model.detect_language和whisper.decode方法仅支持处理30秒时长的单段音频(也就是pad_or_trim处理后固定形状的音频数据),直接传入完整长音频会因为形状不匹配触发断言错误。

正确解决方案:使用transcribe方法处理完整音频

不需要手动修改音频结构,直接用Whisper内置的transcribe方法即可,它会自动处理长音频的分段、语言检测、完整转录流程。修改后的脚本如下:

import whisper

# 加载模型
model = whisper.load_model("large-v2")

# 直接调用transcribe处理完整音频,自动处理分段和语言检测
result = model.transcribe("/content/file.mp3", fp16=False)

# 打印检测到的语言和转录文本
print(f"Detected language: {result['language']}")
print(result['text'])

# 保存转录结果到文件
try:
    with open("output_of_file.txt", "w") as f:
        f.write(result['text'])
        print("Transcription saved to file.")
except Exception as e:
    print(f"Error while saving transcription: {e}")

为什么原来的方法不行?

你原来的代码流程是针对单段30秒音频设计的:

  • pad_or_trim(audio)会把音频裁剪或补零到30秒,符合log_mel_spectrogram、detect_language、decode的输入要求
  • 去掉后,完整音频的长度远超过30秒,生成的梅尔频谱图形状不符合模型断言的要求,因此报错

若需手动处理长音频(不推荐)

如果一定要手动实现长音频处理,需要把音频分割成30秒的片段,逐个处理后拼接结果,示例代码如下(用Whisper自身的工具函数):

import whisper
import numpy as np

model = whisper.load_model("large-v2")
audio = whisper.load_audio("/content/file.mp3")

# 计算每个30秒片段的采样点数量(Whisper默认采样率16000,30秒=480000采样点)
chunk_length = 16000 * 30
chunks = [audio[i:i+chunk_length] for i in range(0, len(audio), chunk_length)]

# 处理每个片段
full_text = ""
detected_lang = None

for chunk in chunks:
    # 补零到30秒(如果最后一段不足)
    chunk = whisper.pad_or_trim(chunk)
    mel = whisper.log_mel_spectrogram(chunk).to(model.device)
    
    # 第一次检测语言
    if not detected_lang:
        _, probs = model.detect_language(mel)
        detected_lang = max(probs, key=probs.get)
        print(f"Detected language: {detected_lang}")
    
    # 解码片段
    options = whisper.DecodingOptions(fp16=False, language=detected_lang)
    result = whisper.decode(model, mel, options)
    full_text += result.text + " "

# 打印并保存完整结果
print(full_text.strip())
try:
    with open("output_of_file.txt", "w") as f:
        f.write(full_text.strip())
        print("Transcription saved to file.")
except Exception as e:
    print(f"Error while saving transcription: {e}")

但强烈建议使用transcribe方法,它内部已经做了更优化的分段、上下文衔接处理,效果比手动拼接更好。

内容的提问来源于stack exchange,提问作者muratowski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 10:57:03