You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何在内存中写入临时WAV文件优化音频处理流程

优化音频处理流程:跳过磁盘IO,内存中处理WAV格式数据

你当前的流程频繁进行磁盘读写,这是性能瓶颈所在。可以通过两种方式优化:在内存中构建WAV格式数据,或者直接将原始音频字节转换为模型需要的numpy数组(更高效)。

方案一:使用内存模拟WAV文件(io.BytesIO)

用io.BytesIO替代真实文件,在内存中完成WAV格式封装,彻底避免磁盘IO操作:

import pyaudio
import crepe
import os
import wave
from scipy.io import wavfile
import time as tm
import io  # 新增内存文件操作模块

# 隐藏TF日志
os.environ['TF_CPP_MIN_LOG_LEVEL'] = '2'

CHUNK = 1024 * 4
WIDTH = 2
CHANNELS = 1
RATE = 16000


try:
    conf_list = []
    looptime_list = []
    print("Recording is starting...")
    p = pyaudio.PyAudio()
    stream = p.open(format=p.get_format_from_width(WIDTH),
                    channels=CHANNELS,
                    rate=RATE,
                    input=True,
                    frames_per_buffer=CHUNK)
    while True:
        start_time = tm.time()
        data = stream.read(CHUNK)

        # 在内存中创建WAV文件
        wav_buffer = io.BytesIO()
        wf = wave.open(wav_buffer, 'wb')
        wf.setnchannels(CHANNELS)
        wf.setsampwidth(WIDTH * 2)
        wf.setframerate(RATE)
        wf.writeframes(data)
        wf.close()
        # 将指针移回缓冲区开头,准备读取
        wav_buffer.seek(0)

        # 从内存缓冲区读取WAV数据
        audio = wavfile.read(wav_buffer)[1]

        time, frequency, confidence, activation = crepe.predict(audio, RATE, model_capacity="tiny", step_size=130, verbose=0)

        if len(time) > 1:
            print("few steps in one chunk")
            for n in range(0, len(time)):
                conf_list.append(confidence[n])
                confidence_mark = "🟥"
                if confidence[n] >= 0.4:
                    confidence_mark = "🟩"
                print(f"{confidence_mark} {round(frequency[n])} Hz | {round(confidence[n], 2)}")
        else:
            conf_list.append(confidence[0])
            confidence_mark = "🟥"
            if confidence[0] >= 0.4:
                confidence_mark = "🟩"
            print(f"{confidence_mark} {round(frequency[0])} Hz | {round(confidence[0], 2)}")


        looptime_list.append(tm.time() - start_time)
except KeyboardInterrupt:
    stream.stop_stream()
    stream.close()
    p.terminate()
    print('Recording stopped')
    av_conf = round(sum(conf_list) / len(conf_list), 2)
    av_lt = round(sum(looptime_list) / len(looptime_list), 5)
    print(f"Average accuracy: {av_conf}\n"
          f"Average time on chunk processing: {av_lt}")

方案二:直接转换原始字节为numpy数组(最优解)

实际上crepe.predict需要的是音频采样数据的numpy数组,不需要经过WAV格式封装。直接把PyAudio读取的字节数据转换成numpy数组,跳过WAV封装步骤,性能提升更明显:

import pyaudio
import crepe
import os
import numpy as np  # 新增numpy用于数据转换
import time as tm

# 隐藏TF日志
os.environ['TF_CPP_MIN_LOG_LEVEL'] = '2'

CHUNK = 1024 * 4
WIDTH = 2
CHANNELS = 1
RATE = 16000
# 根据WIDTH匹配数据类型:WIDTH=2对应16位整数
DTYPE = np.int16 if WIDTH == 2 else np.int32


try:
    conf_list = []
    looptime_list = []
    print("Recording is starting...")
    p = pyaudio.PyAudio()
    stream = p.open(format=p.get_format_from_width(WIDTH),
                    channels=CHANNELS,
                    rate=RATE,
                    input=True,
                    frames_per_buffer=CHUNK)
    while True:
        start_time = tm.time()
        data = stream.read(CHUNK)

        # 直接将字节数据转换为numpy数组
        audio = np.frombuffer(data, dtype=DTYPE)
        # 多通道场景需reshape:audio = audio.reshape(-1, CHANNELS),单通道可省略

        time, frequency, confidence, activation = crepe.predict(audio, RATE, model_capacity="tiny", step_size=130, verbose=0)

        if len(time) > 1:
            print("few steps in one chunk")
            for n in range(0, len(time)):
                conf_list.append(confidence[n])
                confidence_mark = "🟥"
                if confidence[n] >= 0.4:
                    confidence_mark = "🟩"
                print(f"{confidence_mark} {round(frequency[n])} Hz | {round(confidence[n], 2)}")
        else:
            conf_list.append(confidence[0])
            confidence_mark = "🟥"
            if confidence[0] >= 0.4:
                confidence_mark = "🟩"
            print(f"{confidence_mark} {round(frequency[0])} Hz | {round(confidence[0], 2)}")


        looptime_list.append(tm.time() - start_time)
except KeyboardInterrupt:
    stream.stop_stream()
    stream.close()
    p.terminate()
    print('Recording stopped')
    av_conf = round(sum(conf_list) / len(conf_list), 2)
    av_lt = round(sum(looptime_list) / len(looptime_list), 5)
    print(f"Average accuracy: {av_conf}\n"
          f"Average time on chunk processing: {av_lt}")

关键说明

  • 方案一改动最小,完全保留原代码的WAV格式处理逻辑,适合需要兼容WAV格式校验的场景。
  • 方案二是最优选择,直接跳过WAV封装/解封装步骤,减少不必要的计算开销,性能提升最显著。
  • 注意数据类型匹配:PyAudio读取的字节格式必须和numpy的dtype对应,WIDTH=2对应np.int16,WIDTH=4对应np.int32。

内容的提问来源于stack exchange,提问作者Eldar81

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 04:50:24