You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python采样PC音频设备并检测特定输出音频?

Hey there! Let's tackle your two problems step by step, with practical code examples you can adapt for your use case.


1. Sampling Your PC's Audio Output Stream

You mentioned struggling with sounddevice—the key here is targeting the loopback/立体声混音 device that captures audio being sent to your speakers, rather than a microphone input. Here's how to make it work:

First, install the required packages:

pip install sounddevice soundfile librosa numpy

Step 1: Find Your Loopback Device

Run this script to list all audio devices and spot your system's loopback device (it’ll usually have "Loopback", "Stereo Mix", or "Monitor" in the name):

import sounddevice as sd

print("Available audio devices:")
for i, device in enumerate(sd.query_devices()):
    print(f"ID {i}: {device['name']}")
    # Check for loopback/monitor device (naming varies by OS)
    if "loopback" in device['name'].lower() or "stereo mix" in device['name'].lower():
        print(f"→ Found loopback device: ID {i}")

Step 2: Capture Audio in Real-Time

Once you have the loopback device ID, use this code to capture and process audio in chunks (we’ll reuse this for sound detection later):

import sounddevice as sd
import numpy as np
import sys

# Replace with your loopback device ID from the previous step
LOOPBACK_DEVICE_ID = 2
SAMPLING_RATE = 44100
CHUNK_DURATION = 0.5  # Process audio in 0.5-second chunks
CHUNK_SIZE = int(SAMPLING_RATE * CHUNK_DURATION)

def audio_callback(indata, frames, time, status):
    if status:
        print(status, file=sys.stderr)
    # Convert stereo audio to mono for simpler processing
    mono_audio = np.mean(indata, axis=1)
    # Pass the chunk to your detection function
    detect_target_sound(mono_audio)

# Start the audio stream
with sd.InputStream(device=LOOPBACK_DEVICE_ID,
                    samplerate=SAMPLING_RATE,
                    blocksize=CHUNK_SIZE,
                    callback=audio_callback):
    print("Listening to audio output... Press Enter to stop.")
    input()

2. Detecting Short Target Sounds (1-2 Seconds)

Chromaprint’s fpcalc struggles with short audio because it relies on longer temporal patterns for fingerprinting. Instead, we’ll use MFCC (Mel-Frequency Cepstral Coefficients)—a feature that captures the unique timbral characteristics of sound—paired with Dynamic Time Warping (DTW) to handle slight timing differences between the target and captured audio.

Step 1: Preprocess Your Target Audio

First, extract MFCC features from your target sound file:

import librosa

TARGET_AUDIO_PATH = "target_sound.wav"
TARGET_SAMPLING_RATE = 44100

# Load target audio and convert to mono
target_audio, _ = librosa.load(TARGET_AUDIO_PATH, sr=TARGET_SAMPLING_RATE, mono=True)
# Normalize volume to reduce volume-based errors
target_audio = librosa.util.normalize(target_audio)

# Extract MFCC features + delta/delta-delta for temporal context
target_mfcc = librosa.feature.mfcc(y=target_audio, sr=TARGET_SAMPLING_RATE, n_mfcc=13)
target_mfcc_delta = librosa.feature.delta(target_mfcc)
target_mfcc_delta2 = librosa.feature.delta(target_mfcc, order=2)
# Combine all features into one array
target_features = np.concatenate([target_mfcc, target_mfcc_delta, target_mfcc_delta2], axis=0)

Step 2: Real-Time Detection with DTW

Add this function to your stream code to compare captured audio chunks with the target features:

from librosa.sequence import dtw

def detect_target_sound(captured_audio):
    # Normalize captured audio
    captured_audio = librosa.util.normalize(captured_audio)
    
    # Extract features from the captured chunk
    captured_mfcc = librosa.feature.mfcc(y=captured_audio, sr=SAMPLING_RATE, n_mfcc=13)
    captured_mfcc_delta = librosa.feature.delta(captured_mfcc)
    captured_mfcc_delta2 = librosa.feature.delta(captured_mfcc, order=2)
    captured_features = np.concatenate([captured_mfcc, captured_mfcc_delta, captured_mfcc_delta2], axis=0)
    
    # Calculate DTW distance (lower value = more similar)
    try:
        dtw_distance, _, _, _ = dtw(captured_features.T, target_features.T)
        # Adjust this threshold based on your testing (lower = stricter match)
        THRESHOLD = 1000
        if dtw_distance < THRESHOLD:
            print("Target sound detected! Triggering action...")
            # Add your desired action here (e.g., simulate a keypress, run a script)
            # trigger_your_action()
    except Exception:
        # Skip chunks that are too short for valid comparison
        pass

Quick Optimization Tips

  • Tweak the Threshold: Test with different audio scenarios to find a threshold that balances false positives and negatives.
  • Reduce Sampling Rate: If you need faster processing, lower the sampling rate (e.g., 22050 instead of 44100) to reduce feature computation time.
  • Use Mel Spectrograms: For simpler, faster detection, replace MFCCs with a Mel spectrogram (remove the delta features and use librosa.feature.melspectrogram() instead).

内容的提问来源于stack exchange,提问作者RenegadeAndy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:37:36