如何用Python采样PC音频设备并检测特定输出音频?
Hey there! Let's tackle your two problems step by step, with practical code examples you can adapt for your use case.
1. Sampling Your PC's Audio Output Stream
You mentioned struggling with sounddevice—the key here is targeting the loopback/立体声混音 device that captures audio being sent to your speakers, rather than a microphone input. Here's how to make it work:
First, install the required packages:
pip install sounddevice soundfile librosa numpy
Step 1: Find Your Loopback Device
Run this script to list all audio devices and spot your system's loopback device (it’ll usually have "Loopback", "Stereo Mix", or "Monitor" in the name):
import sounddevice as sd print("Available audio devices:") for i, device in enumerate(sd.query_devices()): print(f"ID {i}: {device['name']}") # Check for loopback/monitor device (naming varies by OS) if "loopback" in device['name'].lower() or "stereo mix" in device['name'].lower(): print(f"→ Found loopback device: ID {i}")
Step 2: Capture Audio in Real-Time
Once you have the loopback device ID, use this code to capture and process audio in chunks (we’ll reuse this for sound detection later):
import sounddevice as sd import numpy as np import sys # Replace with your loopback device ID from the previous step LOOPBACK_DEVICE_ID = 2 SAMPLING_RATE = 44100 CHUNK_DURATION = 0.5 # Process audio in 0.5-second chunks CHUNK_SIZE = int(SAMPLING_RATE * CHUNK_DURATION) def audio_callback(indata, frames, time, status): if status: print(status, file=sys.stderr) # Convert stereo audio to mono for simpler processing mono_audio = np.mean(indata, axis=1) # Pass the chunk to your detection function detect_target_sound(mono_audio) # Start the audio stream with sd.InputStream(device=LOOPBACK_DEVICE_ID, samplerate=SAMPLING_RATE, blocksize=CHUNK_SIZE, callback=audio_callback): print("Listening to audio output... Press Enter to stop.") input()
2. Detecting Short Target Sounds (1-2 Seconds)
Chromaprint’s fpcalc struggles with short audio because it relies on longer temporal patterns for fingerprinting. Instead, we’ll use MFCC (Mel-Frequency Cepstral Coefficients)—a feature that captures the unique timbral characteristics of sound—paired with Dynamic Time Warping (DTW) to handle slight timing differences between the target and captured audio.
Step 1: Preprocess Your Target Audio
First, extract MFCC features from your target sound file:
import librosa TARGET_AUDIO_PATH = "target_sound.wav" TARGET_SAMPLING_RATE = 44100 # Load target audio and convert to mono target_audio, _ = librosa.load(TARGET_AUDIO_PATH, sr=TARGET_SAMPLING_RATE, mono=True) # Normalize volume to reduce volume-based errors target_audio = librosa.util.normalize(target_audio) # Extract MFCC features + delta/delta-delta for temporal context target_mfcc = librosa.feature.mfcc(y=target_audio, sr=TARGET_SAMPLING_RATE, n_mfcc=13) target_mfcc_delta = librosa.feature.delta(target_mfcc) target_mfcc_delta2 = librosa.feature.delta(target_mfcc, order=2) # Combine all features into one array target_features = np.concatenate([target_mfcc, target_mfcc_delta, target_mfcc_delta2], axis=0)
Step 2: Real-Time Detection with DTW
Add this function to your stream code to compare captured audio chunks with the target features:
from librosa.sequence import dtw def detect_target_sound(captured_audio): # Normalize captured audio captured_audio = librosa.util.normalize(captured_audio) # Extract features from the captured chunk captured_mfcc = librosa.feature.mfcc(y=captured_audio, sr=SAMPLING_RATE, n_mfcc=13) captured_mfcc_delta = librosa.feature.delta(captured_mfcc) captured_mfcc_delta2 = librosa.feature.delta(captured_mfcc, order=2) captured_features = np.concatenate([captured_mfcc, captured_mfcc_delta, captured_mfcc_delta2], axis=0) # Calculate DTW distance (lower value = more similar) try: dtw_distance, _, _, _ = dtw(captured_features.T, target_features.T) # Adjust this threshold based on your testing (lower = stricter match) THRESHOLD = 1000 if dtw_distance < THRESHOLD: print("Target sound detected! Triggering action...") # Add your desired action here (e.g., simulate a keypress, run a script) # trigger_your_action() except Exception: # Skip chunks that are too short for valid comparison pass
Quick Optimization Tips
- Tweak the Threshold: Test with different audio scenarios to find a threshold that balances false positives and negatives.
- Reduce Sampling Rate: If you need faster processing, lower the sampling rate (e.g., 22050 instead of 44100) to reduce feature computation time.
- Use Mel Spectrograms: For simpler, faster detection, replace MFCCs with a Mel spectrogram (remove the delta features and use
librosa.feature.melspectrogram()instead).
内容的提问来源于stack exchange,提问作者RenegadeAndy

