You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将float32格式麦克风音频直接传入Whisper的transcribe()?

直接将麦克风捕获的float32音频数组传入Whisper的解决方案

问题描述

想用sounddevice的rec()函数从麦克风捕获音频,以float32格式存储后直接传入Whisper的transcribe()函数,跳过保存文件再读取的步骤。但运行代码时触发内存溢出错误,提示尝试分配847GB内存,明显存在异常。

尝试的代码

import sounddevice as sd
import numpy as np
import whisper

duration = 15
samplerate = 44100

frames = duration * samplerate

recording = sd.rec(frames, blocking=True, dtype='float32')

model = whisper.load_model("tiny")
rec_array = np.array(recording,dtype=np.float32)
result = model.transcribe(recording,word_timestamps=True, fp16=False)
text = result["text"].strip()
print(text)

报错信息

RuntimeError: [enforce fail at alloc_cpu.cpp:80] data. DefaultCPUAllocator: not enough memory: you tried to allocate 846721764000 bytes.

完整报错回溯

Traceback (most recent call last):
  File "d:\Seshrut\Error-505!!\projects\learn whisper\soundbreak.py", line 34, in <module>
    result = model.transcribe(recording,word_timestamps=True)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\seshr\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\whisper\transcribe.py", line 133, in transcribe
    mel = log_mel_spectrogram(audio, model.dims.n_mels, padding=N_SAMPLES)
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\seshr\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\whisper\audio.py", line 146, in log_mel_spectrogram
    audio = F.pad(audio, (0, padding))
            ^^^^^^^^^^^^^^^^^^^^^^^^^^
RuntimeError: [enforce fail at alloc_cpu.cpp:80] data. DefaultCPUAllocator: not enough memory: you tried to allocate 846721764000 bytes.

问题原因

  1. 声道格式不匹配:Whisper的transcribe()要求输入单声道的1维数组,但sounddevice默认捕获的是立体声(2维数组,形状为(帧数, 2))。传入2维数组后,Whisper会错误地将其解析成长度异常的单声道数据,导致内存分配请求爆炸。
  2. 采样率不匹配:Whisper默认期望16000Hz的音频输入,当前代码用44100Hz捕获,会导致转录时的速度、时长匹配错误。

修正后的代码

import sounddevice as sd
import numpy as np
import whisper

duration = 15
target_samplerate = 16000  # Whisper标准采样率
original_samplerate = 44100

# 捕获立体声音频
recording = sd.rec(
    int(duration * original_samplerate),
    samplerate=original_samplerate,
    blocking=True,
    dtype='float32',
    channels=2
)

# 转单声道:取两个声道的平均值
mono_audio = np.mean(recording, axis=1)

# 转换采样率到16000Hz
mono_audio_resampled = sd.resample(mono_audio, original_samplerate, target_samplerate)

# 加载模型并转录
model = whisper.load_model("tiny")
result = model.transcribe(mono_audio_resampled, word_timestamps=True, fp16=False)
text = result["text"].strip()
print(text)

关键修正点

  • 用np.mean(recording, axis=1)将2维立体声数组转为1维单声道数组
  • 通过sd.resample()把采样率从44100Hz转换为Whisper要求的16000Hz
  • 确保传入transcribe()的是符合要求的1维float32数组

内容的提问来源于stack exchange,提问作者Seshrut

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 08:05:27