You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何关闭OpenAI Whisper归一化并转录超30秒音频?

问题描述

我是OpenAI Whisper和Python的完全新手,需要Whisper生成包含填充词(ah、mh、mhm、uh、oh等)的原始转录文本。已知设置normalize=False可关闭归一化,但当前代码仅能处理30秒以内的音频。尝试结合transformers的pipeline处理长音频,但不知道如何加载本地MP3文件,也不清楚在哪里设置normalize=False。

当前使用的代码:

from transformers import WhisperProcessor, WhisperForConditionalGeneration
import librosa

speech, _ = librosa.load("myaudio.mp3", sr=16000, mono=True)

processor = WhisperProcessor.from_pretrained("openai/whisper-large")
model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-large")

model.config.forced_decoder_ids = processor.get_decoder_prompt_ids(language = "de", task = "transcribe")
input_features = processor(speech, return_tensors="pt", sampling_rate=16000).input_features 
predicted_ids = model.generate(input_features)
transcription = processor.batch_decode(predicted_ids, skip_special_tokens = True, normalize = False)

print(transcription)

参考的长音频处理代码:

import torch
from transformers import pipeline
from datasets import load_dataset
device = "cuda:0" if torch.cuda.is_available() else "cpu"
pipe = pipeline(
  "automatic-speech-recognition",
  model="openai/whisper-base",
  chunk_length_s=30,
  device=device,
)
ds = load_dataset("hf-internal-testing/librispeech_asr_dummy", "clean", split="validation")
sample = ds[0]["audio"]

prediction = pipe(sample.copy())["text"]

# 也可以返回预测的时间戳
prediction = pipe(sample, return_timestamps=True)["chunks"]

解决方案

1. 加载本地MP3文件

transformers的ASR pipeline支持两种便捷的本地音频加载方式:

  • 直接传入本地MP3文件路径字符串,pipeline会自动完成加载、采样率转换(转为Whisper要求的16000Hz)
  • 用librosa/soundfile读取音频后,构造包含array(音频数据数组)和sampling_rate(采样率)的字典传入

最简便的方式是直接传文件路径:

audio_path = "myaudio.mp3"

2. 设置normalize=False保留填充词

在调用pipeline时,通过decode_kwargs参数将normalize=False传递给处理器的解码步骤,即可保留原始填充词。同时需要传入德语的强制解码提示词,确保模型正确识别德语音频。

3. 完整可运行代码

import torch
from transformers import pipeline, WhisperProcessor

device = "cuda:0" if torch.cuda.is_available() else "cpu"

# 初始化处理器,获取德语转录的强制解码提示词
processor = WhisperProcessor.from_pretrained("openai/whisper-large")
forced_decoder_ids = processor.get_decoder_prompt_ids(language="de", task="transcribe")

# 初始化长音频处理pipeline
pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large",
    chunk_length_s=30,  # 按30秒分片处理长音频
    device=device,
    generate_kwargs={"forced_decoder_ids": forced_decoder_ids},  # 指定德语转录
)

# 加载本地MP3并生成带填充词的转录文本
full_transcript = pipe(
    "myaudio.mp3",
    decode_kwargs={"skip_special_tokens": True, "normalize": False}
)
print("完整转录文本:", full_transcript["text"])

# 如需时间戳,开启return_timestamps参数
transcript_with_timestamps = pipe(
    "myaudio.mp3",
    return_timestamps=True,
    decode_kwargs={"skip_special_tokens": True, "normalize": False}
)
print("带时间戳的转录片段:", transcript_with_timestamps["chunks"])

关键说明

  • chunk_length_s=30:自动将长音频切分为30秒分片处理,解决原代码仅支持短音频的问题
  • generate_kwargs:传入德语强制解码提示词,避免模型误识别语言
  • decode_kwargs={"normalize": False}:关闭文本归一化,保留ah、mh等原始填充词
  • 直接传文件路径时,pipeline会自动处理音频格式、采样率转换,无需手动预处理

内容的提问来源于stack exchange,提问作者Psychic Birdy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 20:42:56