如何关闭OpenAI Whisper归一化并转录超30秒音频?
问题描述
我是OpenAI Whisper和Python的完全新手,需要Whisper生成包含填充词(ah、mh、mhm、uh、oh等)的原始转录文本。已知设置normalize=False可关闭归一化,但当前代码仅能处理30秒以内的音频。尝试结合transformers的pipeline处理长音频,但不知道如何加载本地MP3文件,也不清楚在哪里设置normalize=False。
当前使用的代码:
from transformers import WhisperProcessor, WhisperForConditionalGeneration import librosa speech, _ = librosa.load("myaudio.mp3", sr=16000, mono=True) processor = WhisperProcessor.from_pretrained("openai/whisper-large") model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-large") model.config.forced_decoder_ids = processor.get_decoder_prompt_ids(language = "de", task = "transcribe") input_features = processor(speech, return_tensors="pt", sampling_rate=16000).input_features predicted_ids = model.generate(input_features) transcription = processor.batch_decode(predicted_ids, skip_special_tokens = True, normalize = False) print(transcription)
参考的长音频处理代码:
import torch from transformers import pipeline from datasets import load_dataset device = "cuda:0" if torch.cuda.is_available() else "cpu" pipe = pipeline( "automatic-speech-recognition", model="openai/whisper-base", chunk_length_s=30, device=device, ) ds = load_dataset("hf-internal-testing/librispeech_asr_dummy", "clean", split="validation") sample = ds[0]["audio"] prediction = pipe(sample.copy())["text"] # 也可以返回预测的时间戳 prediction = pipe(sample, return_timestamps=True)["chunks"]
解决方案
1. 加载本地MP3文件
transformers的ASR pipeline支持两种便捷的本地音频加载方式:
- 直接传入本地MP3文件路径字符串,pipeline会自动完成加载、采样率转换(转为Whisper要求的16000Hz)
- 用librosa/soundfile读取音频后,构造包含
array(音频数据数组)和sampling_rate(采样率)的字典传入
最简便的方式是直接传文件路径:
audio_path = "myaudio.mp3"
2. 设置normalize=False保留填充词
在调用pipeline时,通过decode_kwargs参数将normalize=False传递给处理器的解码步骤,即可保留原始填充词。同时需要传入德语的强制解码提示词,确保模型正确识别德语音频。
3. 完整可运行代码
import torch from transformers import pipeline, WhisperProcessor device = "cuda:0" if torch.cuda.is_available() else "cpu" # 初始化处理器,获取德语转录的强制解码提示词 processor = WhisperProcessor.from_pretrained("openai/whisper-large") forced_decoder_ids = processor.get_decoder_prompt_ids(language="de", task="transcribe") # 初始化长音频处理pipeline pipe = pipeline( "automatic-speech-recognition", model="openai/whisper-large", chunk_length_s=30, # 按30秒分片处理长音频 device=device, generate_kwargs={"forced_decoder_ids": forced_decoder_ids}, # 指定德语转录 ) # 加载本地MP3并生成带填充词的转录文本 full_transcript = pipe( "myaudio.mp3", decode_kwargs={"skip_special_tokens": True, "normalize": False} ) print("完整转录文本:", full_transcript["text"]) # 如需时间戳,开启return_timestamps参数 transcript_with_timestamps = pipe( "myaudio.mp3", return_timestamps=True, decode_kwargs={"skip_special_tokens": True, "normalize": False} ) print("带时间戳的转录片段:", transcript_with_timestamps["chunks"])
关键说明
chunk_length_s=30:自动将长音频切分为30秒分片处理,解决原代码仅支持短音频的问题generate_kwargs:传入德语强制解码提示词,避免模型误识别语言decode_kwargs={"normalize": False}:关闭文本归一化,保留ah、mh等原始填充词- 直接传文件路径时,pipeline会自动处理音频格式、采样率转换,无需手动预处理
内容的提问来源于stack exchange,提问作者Psychic Birdy
相关产品推荐
相关产品推荐

