Python中如何在内存中写入临时WAV文件优化音频处理流程
优化音频处理流程:跳过磁盘IO,内存中处理WAV格式数据
你当前的流程频繁进行磁盘读写,这是性能瓶颈所在。可以通过两种方式优化:在内存中构建WAV格式数据,或者直接将原始音频字节转换为模型需要的numpy数组(更高效)。
方案一:使用内存模拟WAV文件(io.BytesIO)
用io.BytesIO替代真实文件,在内存中完成WAV格式封装,彻底避免磁盘IO操作:
import pyaudio import crepe import os import wave from scipy.io import wavfile import time as tm import io # 新增内存文件操作模块 # 隐藏TF日志 os.environ['TF_CPP_MIN_LOG_LEVEL'] = '2' CHUNK = 1024 * 4 WIDTH = 2 CHANNELS = 1 RATE = 16000 try: conf_list = [] looptime_list = [] print("Recording is starting...") p = pyaudio.PyAudio() stream = p.open(format=p.get_format_from_width(WIDTH), channels=CHANNELS, rate=RATE, input=True, frames_per_buffer=CHUNK) while True: start_time = tm.time() data = stream.read(CHUNK) # 在内存中创建WAV文件 wav_buffer = io.BytesIO() wf = wave.open(wav_buffer, 'wb') wf.setnchannels(CHANNELS) wf.setsampwidth(WIDTH * 2) wf.setframerate(RATE) wf.writeframes(data) wf.close() # 将指针移回缓冲区开头,准备读取 wav_buffer.seek(0) # 从内存缓冲区读取WAV数据 audio = wavfile.read(wav_buffer)[1] time, frequency, confidence, activation = crepe.predict(audio, RATE, model_capacity="tiny", step_size=130, verbose=0) if len(time) > 1: print("few steps in one chunk") for n in range(0, len(time)): conf_list.append(confidence[n]) confidence_mark = "🟥" if confidence[n] >= 0.4: confidence_mark = "🟩" print(f"{confidence_mark} {round(frequency[n])} Hz | {round(confidence[n], 2)}") else: conf_list.append(confidence[0]) confidence_mark = "🟥" if confidence[0] >= 0.4: confidence_mark = "🟩" print(f"{confidence_mark} {round(frequency[0])} Hz | {round(confidence[0], 2)}") looptime_list.append(tm.time() - start_time) except KeyboardInterrupt: stream.stop_stream() stream.close() p.terminate() print('Recording stopped') av_conf = round(sum(conf_list) / len(conf_list), 2) av_lt = round(sum(looptime_list) / len(looptime_list), 5) print(f"Average accuracy: {av_conf}\n" f"Average time on chunk processing: {av_lt}")
方案二:直接转换原始字节为numpy数组(最优解)
实际上crepe.predict需要的是音频采样数据的numpy数组,不需要经过WAV格式封装。直接把PyAudio读取的字节数据转换成numpy数组,跳过WAV封装步骤,性能提升更明显:
import pyaudio import crepe import os import numpy as np # 新增numpy用于数据转换 import time as tm # 隐藏TF日志 os.environ['TF_CPP_MIN_LOG_LEVEL'] = '2' CHUNK = 1024 * 4 WIDTH = 2 CHANNELS = 1 RATE = 16000 # 根据WIDTH匹配数据类型:WIDTH=2对应16位整数 DTYPE = np.int16 if WIDTH == 2 else np.int32 try: conf_list = [] looptime_list = [] print("Recording is starting...") p = pyaudio.PyAudio() stream = p.open(format=p.get_format_from_width(WIDTH), channels=CHANNELS, rate=RATE, input=True, frames_per_buffer=CHUNK) while True: start_time = tm.time() data = stream.read(CHUNK) # 直接将字节数据转换为numpy数组 audio = np.frombuffer(data, dtype=DTYPE) # 多通道场景需reshape:audio = audio.reshape(-1, CHANNELS),单通道可省略 time, frequency, confidence, activation = crepe.predict(audio, RATE, model_capacity="tiny", step_size=130, verbose=0) if len(time) > 1: print("few steps in one chunk") for n in range(0, len(time)): conf_list.append(confidence[n]) confidence_mark = "🟥" if confidence[n] >= 0.4: confidence_mark = "🟩" print(f"{confidence_mark} {round(frequency[n])} Hz | {round(confidence[n], 2)}") else: conf_list.append(confidence[0]) confidence_mark = "🟥" if confidence[0] >= 0.4: confidence_mark = "🟩" print(f"{confidence_mark} {round(frequency[0])} Hz | {round(confidence[0], 2)}") looptime_list.append(tm.time() - start_time) except KeyboardInterrupt: stream.stop_stream() stream.close() p.terminate() print('Recording stopped') av_conf = round(sum(conf_list) / len(conf_list), 2) av_lt = round(sum(looptime_list) / len(looptime_list), 5) print(f"Average accuracy: {av_conf}\n" f"Average time on chunk processing: {av_lt}")
关键说明
- 方案一改动最小,完全保留原代码的WAV格式处理逻辑,适合需要兼容WAV格式校验的场景。
- 方案二是最优选择,直接跳过WAV封装/解封装步骤,减少不必要的计算开销,性能提升最显著。
- 注意数据类型匹配:PyAudio读取的字节格式必须和numpy的dtype对应,WIDTH=2对应
np.int16,WIDTH=4对应np.int32。
内容的提问来源于stack exchange,提问作者Eldar81
相关产品推荐
相关产品推荐

