如何将float32格式麦克风音频直接传入Whisper的transcribe()?
直接将麦克风捕获的float32音频数组传入Whisper的解决方案
问题描述
想用sounddevice的rec()函数从麦克风捕获音频,以float32格式存储后直接传入Whisper的transcribe()函数,跳过保存文件再读取的步骤。但运行代码时触发内存溢出错误,提示尝试分配847GB内存,明显存在异常。
尝试的代码
import sounddevice as sd import numpy as np import whisper duration = 15 samplerate = 44100 frames = duration * samplerate recording = sd.rec(frames, blocking=True, dtype='float32') model = whisper.load_model("tiny") rec_array = np.array(recording,dtype=np.float32) result = model.transcribe(recording,word_timestamps=True, fp16=False) text = result["text"].strip() print(text)
报错信息
RuntimeError: [enforce fail at alloc_cpu.cpp:80] data. DefaultCPUAllocator: not enough memory: you tried to allocate 846721764000 bytes.
完整报错回溯
Traceback (most recent call last): File "d:\Seshrut\Error-505!!\projects\learn whisper\soundbreak.py", line 34, in <module> result = model.transcribe(recording,word_timestamps=True) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\seshr\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\whisper\transcribe.py", line 133, in transcribe mel = log_mel_spectrogram(audio, model.dims.n_mels, padding=N_SAMPLES) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\seshr\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\whisper\audio.py", line 146, in log_mel_spectrogram audio = F.pad(audio, (0, padding)) ^^^^^^^^^^^^^^^^^^^^^^^^^^ RuntimeError: [enforce fail at alloc_cpu.cpp:80] data. DefaultCPUAllocator: not enough memory: you tried to allocate 846721764000 bytes.
问题原因
- 声道格式不匹配:Whisper的
transcribe()要求输入单声道的1维数组,但sounddevice默认捕获的是立体声(2维数组,形状为(帧数, 2))。传入2维数组后,Whisper会错误地将其解析成长度异常的单声道数据,导致内存分配请求爆炸。 - 采样率不匹配:Whisper默认期望16000Hz的音频输入,当前代码用44100Hz捕获,会导致转录时的速度、时长匹配错误。
修正后的代码
import sounddevice as sd import numpy as np import whisper duration = 15 target_samplerate = 16000 # Whisper标准采样率 original_samplerate = 44100 # 捕获立体声音频 recording = sd.rec( int(duration * original_samplerate), samplerate=original_samplerate, blocking=True, dtype='float32', channels=2 ) # 转单声道:取两个声道的平均值 mono_audio = np.mean(recording, axis=1) # 转换采样率到16000Hz mono_audio_resampled = sd.resample(mono_audio, original_samplerate, target_samplerate) # 加载模型并转录 model = whisper.load_model("tiny") result = model.transcribe(mono_audio_resampled, word_timestamps=True, fp16=False) text = result["text"].strip() print(text)
关键修正点
- 用
np.mean(recording, axis=1)将2维立体声数组转为1维单声道数组 - 通过
sd.resample()把采样率从44100Hz转换为Whisper要求的16000Hz - 确保传入
transcribe()的是符合要求的1维float32数组
内容的提问来源于stack exchange,提问作者Seshrut
相关产品推荐
相关产品推荐

