You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实时音频音符识别报错:OSError: [Error Not input stream] -9975

问题解决:OSError: [Error Not input stream] -9975

错误根源

核心问题出在音频流的配置逻辑上:

  • 你打开音频流时设置了input=False,但后续调用了stream.read(1024)——read()方法仅支持输入流(input=True),输出流无法读取数据,直接触发该错误。
  • 代码注释里“Set input=False to record from speaker”是错误逻辑,采集音频(无论麦克风还是系统内放)都需要将流设置为输入流。

修复步骤

1. 修正音频流的输入输出配置

要采集音频,必须将input设为True,无需播放音频时可将output设为False:

stream = p.open(format=pyaudio.paInt16,
                channels=1,
                rate=44100,
                output=False,  # 不需要播放则关闭输出
                input=True)   # 采集音频必须开启输入

2. 修复FFT频率计算错误

原代码的fftfreq调用缺少采样率参数,导致计算出的是归一化频率而非实际Hz值,同时FFT结果包含正负频率,只需保留正频率部分:

# 补充采样率参数,计算实际频率
frequencies = fftfreq(len(audio_data), d=1/44100)
# 过滤出正频率
positive_mask = frequencies > 0
positive_frequencies = frequencies[positive_mask]
positive_fft_abs = np.abs(fft)[positive_mask]

3. 完善频率转音符函数

原determine_note函数未实现,补充标准的频率转音符逻辑:

def determine_note(frequency):
    if frequency <= 20:  # 过滤低频噪音
        return ""
    A4 = 440  # 标准A4音频率
    half_steps = 12 * math.log2(frequency / A4)
    rounded_steps = round(half_steps)
    note_names = ["C", "C#", "D", "D#", "E", "F", "F#", "G", "G#", "A", "A#", "B"]
    note_index = (rounded_steps + 9) % 12  # A4对应索引9,偏移后取模
    octave = 4 + (rounded_steps + 9) // 12
    return f"{note_names[note_index]}{octave}"

完整修复代码

import pyaudio
import numpy as np
from scipy.fftpack import fftfreq
from scipy.signal import find_peaks
import math

def process_audio(audio_data, sample_rate=44100):
    fft = np.fft.fft(audio_data)
    frequencies = fftfreq(len(audio_data), d=1/sample_rate)
    
    # 只保留正频率部分
    positive_mask = frequencies > 0
    positive_frequencies = frequencies[positive_mask]
    positive_fft_abs = np.abs(fft)[positive_mask]

    peak_indices, _ = find_peaks(positive_fft_abs)
    peak_frequencies = positive_frequencies[peak_indices]
    peak_magnitudes = positive_fft_abs[peak_indices]

    # 阈值过滤小幅度峰值
    if len(peak_magnitudes) == 0:
        return []
    threshold = 0.2 * np.max(peak_magnitudes)
    valid_peaks = peak_frequencies[peak_magnitudes > threshold]

    # 转换为音符并过滤噪音
    detectable_notes = [determine_note(f) for f in valid_peaks if f > 20]
    return detectable_notes


def determine_note(frequency):
    A4 = 440
    half_steps = 12 * math.log2(frequency / A4)
    rounded_steps = round(half_steps)
    note_names = ["C", "C#", "D", "D#", "E", "F", "F#", "G", "G#", "A", "A#", "B"]
    note_index = (rounded_steps + 9) % 12
    octave = 4 + (rounded_steps + 9) // 12
    return f"{note_names[note_index]}{octave}"

if __name__ == "__main__":
    p = pyaudio.PyAudio()

    # 打开输入流采集音频
    stream = p.open(format=pyaudio.paInt16,
                    channels=1,
                    rate=44100,
                    output=False,
                    input=True)

    try:
        while True:
            data = stream.read(1024)
            audio_data = np.frombuffer(data, dtype=np.int16)

            detected_notes = process_audio(audio_data)
            if detected_notes:
                print(f"识别到音符: {', '.join(detected_notes)}")
            else:
                print("未识别到音符")
    except KeyboardInterrupt:
        print("\n程序终止")
    finally:
        stream.stop_stream()
        stream.close()
        p.terminate()

额外提示

如果需要采集电脑播放的音频(而非麦克风输入),需将系统默认输入设备设置为“立体声混音”,或通过以下代码查看设备索引后指定:

p = pyaudio.PyAudio()
for i in range(p.get_device_count()):
    info = p.get_device_info_by_index(i)
    print(f"设备索引 {i}: {info['name']}")

找到立体声混音的索引后,在p.open()中添加参数:input_device_index=你的设备索引

内容的提问来源于stack exchange,提问作者Kislok

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 00:32:04