Python实时音频音符识别报错:OSError: [Error Not input stream] -9975
问题解决:OSError: [Error Not input stream] -9975
错误根源
核心问题出在音频流的配置逻辑上:
- 你打开音频流时设置了
input=False,但后续调用了stream.read(1024)——read()方法仅支持输入流(input=True),输出流无法读取数据,直接触发该错误。 - 代码注释里“Set input=False to record from speaker”是错误逻辑,采集音频(无论麦克风还是系统内放)都需要将流设置为输入流。
修复步骤
1. 修正音频流的输入输出配置
要采集音频,必须将input设为True,无需播放音频时可将output设为False:
stream = p.open(format=pyaudio.paInt16, channels=1, rate=44100, output=False, # 不需要播放则关闭输出 input=True) # 采集音频必须开启输入
2. 修复FFT频率计算错误
原代码的fftfreq调用缺少采样率参数,导致计算出的是归一化频率而非实际Hz值,同时FFT结果包含正负频率,只需保留正频率部分:
# 补充采样率参数,计算实际频率 frequencies = fftfreq(len(audio_data), d=1/44100) # 过滤出正频率 positive_mask = frequencies > 0 positive_frequencies = frequencies[positive_mask] positive_fft_abs = np.abs(fft)[positive_mask]
3. 完善频率转音符函数
原determine_note函数未实现,补充标准的频率转音符逻辑:
def determine_note(frequency): if frequency <= 20: # 过滤低频噪音 return "" A4 = 440 # 标准A4音频率 half_steps = 12 * math.log2(frequency / A4) rounded_steps = round(half_steps) note_names = ["C", "C#", "D", "D#", "E", "F", "F#", "G", "G#", "A", "A#", "B"] note_index = (rounded_steps + 9) % 12 # A4对应索引9,偏移后取模 octave = 4 + (rounded_steps + 9) // 12 return f"{note_names[note_index]}{octave}"
完整修复代码
import pyaudio import numpy as np from scipy.fftpack import fftfreq from scipy.signal import find_peaks import math def process_audio(audio_data, sample_rate=44100): fft = np.fft.fft(audio_data) frequencies = fftfreq(len(audio_data), d=1/sample_rate) # 只保留正频率部分 positive_mask = frequencies > 0 positive_frequencies = frequencies[positive_mask] positive_fft_abs = np.abs(fft)[positive_mask] peak_indices, _ = find_peaks(positive_fft_abs) peak_frequencies = positive_frequencies[peak_indices] peak_magnitudes = positive_fft_abs[peak_indices] # 阈值过滤小幅度峰值 if len(peak_magnitudes) == 0: return [] threshold = 0.2 * np.max(peak_magnitudes) valid_peaks = peak_frequencies[peak_magnitudes > threshold] # 转换为音符并过滤噪音 detectable_notes = [determine_note(f) for f in valid_peaks if f > 20] return detectable_notes def determine_note(frequency): A4 = 440 half_steps = 12 * math.log2(frequency / A4) rounded_steps = round(half_steps) note_names = ["C", "C#", "D", "D#", "E", "F", "F#", "G", "G#", "A", "A#", "B"] note_index = (rounded_steps + 9) % 12 octave = 4 + (rounded_steps + 9) // 12 return f"{note_names[note_index]}{octave}" if __name__ == "__main__": p = pyaudio.PyAudio() # 打开输入流采集音频 stream = p.open(format=pyaudio.paInt16, channels=1, rate=44100, output=False, input=True) try: while True: data = stream.read(1024) audio_data = np.frombuffer(data, dtype=np.int16) detected_notes = process_audio(audio_data) if detected_notes: print(f"识别到音符: {', '.join(detected_notes)}") else: print("未识别到音符") except KeyboardInterrupt: print("\n程序终止") finally: stream.stop_stream() stream.close() p.terminate()
额外提示
如果需要采集电脑播放的音频(而非麦克风输入),需将系统默认输入设备设置为“立体声混音”,或通过以下代码查看设备索引后指定:
p = pyaudio.PyAudio() for i in range(p.get_device_count()): info = p.get_device_info_by_index(i) print(f"设备索引 {i}: {info['name']}")
找到立体声混音的索引后,在p.open()中添加参数:input_device_index=你的设备索引
内容的提问来源于stack exchange,提问作者Kislok
相关产品推荐
相关产品推荐

