You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Librosa的特定报警音频检测脚本误触发问题求助

基于Librosa的特定报警音频检测脚本误触发问题求助

大家好,我目前在开发一个音频触发的自动点击脚本,遇到了精准度的问题,希望能得到大家的帮助。

我的核心需求是:当电脑播放我指定的3秒报警音频(ALARM_BIRD.mp3)时,脚本立即停止监听并执行3次指定坐标的鼠标点击。但现在脚本的表现完全失控——它会随机触发:有时候是其他无关声音,甚至电脑处于安静状态时,也会莫名其妙地执行点击操作,完全没法精准识别我的目标报警音。

我的实现逻辑

  • 用librosa加载参考报警音频,提取MFCC特征的均值作为识别基准
  • 用sounddevice实时监听电脑的输入音频,对每一段实时音频同样提取MFCC特征均值
  • 用欧氏距离计算实时特征与基准特征的差距,当差距小于设定的THRESHOLD阈值时,触发3次鼠标点击并终止程序

尝试过的调整(均无效)

  • 修改librosa.load中的采样率参数(比如调整为44100或其他数值)
  • 多次修改THRESHOLD的阈值大小
  • 确认参考音频是我需要的3秒特定报警音

脚本运行时会实时打印当前音频与参考音频的特征距离,理论上只有当距离处于0到THRESHOLD区间内时才会触发,但实际情况是触发时机完全随机,和我的目标报警音播放完全不匹配。我实在找不到问题所在了,希望有人能帮我调整脚本,让它只在目标报警音播放时才触发点击。

以下是我的完整代码:

import sounddevice as sd
import numpy as np
import librosa
import time
import pyautogui
from scipy.spatial.distance import euclidean

# Configuration
REFERENCE_AUDIO_FILE = "ALARM_BIRD.mp3"  # Path to the reference audio file
DURATION = 120                           # Monitoring time in seconds
CLICK_INTERVAL = 0.1                     # Interval between clicks
THRESHOLD = 40                           # Similarity threshold for detection

# Load the reference audio file
def load_reference_audio(file):
    print("Loading reference audio...")
    y, sr = librosa.load(file, sr=44100)
    mfcc = librosa.feature.mfcc(y=y, sr=sr)
    return mfcc.mean(axis=1)

# Monitor audio in real-time
def monitor_audio(reference_features):
    print("Monitoring audio...")

    def callback(indata, frames, time, status):
        y = indata[:, 0]  # Use the first audio channel
        mfcc = librosa.feature.mfcc(y=y, sr=44100)
        current_features = mfcc.mean(axis=1)

        # Calculate the distance between the current sound and the reference sound
        distance = euclidean(reference_features, current_features)
        print(f"Distance to reference sound: {distance}")

        if distance < THRESHOLD:  # If the distance is below the threshold
            print("Sound detected! Performing three clicks.")
            perform_clicks()
            exit()  # Stop the program after detection

    with sd.InputStream(callback=callback):
        time.sleep(DURATION)

# Perform three mouse clicks
def perform_clicks():
    for _ in range(3):
        pyautogui.click(-1519, 342)  # Change (-1519, 342) to the desired coordinates
        print("Click performed.")
        time.sleep(CLICK_INTERVAL)

# Run the program
if __name__ == "__main__":
    reference_features = load_reference_audio(REFERENCE_AUDIO_FILE)
    monitor_audio(reference_features)

备注:内容来源于stack exchange,提问作者Fran Coronas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 12:23:04