You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用Python Dejavu库检测单个音频文件中的多首歌曲?

Detecting Multiple Songs in a Single MP3 with Dejavu

Great question! I’ve dealt with this exact limitation of Dejavu before—out of the box, it’s optimized to return the first matching track it finds, which doesn’t work for multi-song files. Here are three practical solutions to modify its behavior and identify all tracks in your MP3:

1. Split the Audio on Silence & Recognize Chunks

The easiest approach is to split the MP3 into individual song chunks using silence detection, then run Dejavu on each chunk. This mimics how you’d manually split a mixed MP3, and plays nicely with Dejavu’s existing workflow.

Use a library like pydub to handle the audio splitting:

from pydub import AudioSegment
from pydub.silence import split_on_silence
import dejavu
import os

# Initialize Dejavu with your database config
config = {
    "database": {
        "host": "localhost",
        "user": "root",
        "password": "",
        "database": "dejavu"
    }
}
djv = dejavu.Dejavu(config)

# Load your multi-song MP3
audio = AudioSegment.from_mp3("mixed_songs.mp3")

# Split audio into chunks based on silence (tune these values to your file)
chunks = split_on_silence(
    audio,
    min_silence_len=600,    # Minimum silence duration to trigger a split (ms)
    silence_thresh=-35,     # Volume threshold for "silence" (dB)
    keep_silence=300        # Keep a small amount of silence around chunks
)

# Recognize each chunk and collect unique results
detected_songs = []
for idx, chunk in enumerate(chunks):
    temp_file = f"temp_chunk_{idx}.wav"
    chunk.export(temp_file, format="wav")
    
    match = djv.recognize("file", temp_file)
    if match and match["song_name"] not in [s["song_name"] for s in detected_songs]:
        detected_songs.append(match)
    
    # Clean up temporary files
    os.remove(temp_file)

print("Detected songs in order:")
for song in detected_songs:
    print(f"- {song['song_name']} (confidence: {song['confidence']})")

2. Modify Dejavu’s Core Recognition Logic

If you want to avoid splitting audio, you can tweak Dejavu’s source code to scan the entire file and collect all matches instead of stopping at the first one.

Navigate to the recognize.py file in your Dejavu installation (usually in dejavu/recognize.py) and modify the recognize_file method:

Original code (abbreviated):

# ... after calculating matches ...
for song_id, offset in sorted(matches.items(), key=lambda x: x[1]):
    song = self.db.get_song_by_id(song_id)
    return {
        'song_id': song_id,
        'song_name': song['song_name'],
        'confidence': len(matches[song_id]),
        'offset': offset,
        'offset_seconds': offset / self.FS,
        'file_sha1': sha1
    }

Modified code to return all matches:

# ... after calculating matches ...
all_detected_songs = []
for song_id, offsets in matches.items():
    song = self.db.get_song_by_id(song_id)
    # Calculate earliest offset for ordering
    earliest_offset = min(offsets)
    all_detected_songs.append({
        'song_id': song_id,
        'song_name': song['song_name'],
        'confidence': len(offsets),
        'earliest_offset': earliest_offset,
        'earliest_offset_seconds': earliest_offset / self.FS,
        'file_sha1': sha1
    })

# Sort songs by their earliest appearance in the file
all_detected_songs.sort(key=lambda x: x['earliest_offset_seconds'])
return all_detected_songs

Now when you call djv.recognize("file", "mixed_songs.mp3"), it will return a list of all matched songs ordered by their start time.

3. Sliding Window Scan

For more control (e.g., if silence detection isn’t reliable), use a sliding window approach: scan the MP3 in overlapping segments, run Dejavu on each segment, and deduplicate results.

import dejavu
from pydub import AudioSegment
import os

config = {
    "database": {
        "host": "localhost",
        "user": "root",
        "password": "",
        "database": "dejavu"
    }
}
djv = dejavu.Dejavu(config)

audio = AudioSegment.from_mp3("mixed_songs.mp3")
window_length = 12 * 1000  # 12-second scan window
step_length = 4 * 1000     # Move window forward by 4 seconds each time
total_length = len(audio)

# Track detected songs and their first occurrence time
detected_songs = {}

for start_ms in range(0, total_length - window_length, step_length):
    end_ms = start_ms + window_length
    window = audio[start_ms:end_ms]
    
    temp_file = "temp_window.wav"
    window.export(temp_file, format="wav")
    
    match = djv.recognize("file", temp_file)
    if match:
        song_name = match["song_name"]
        # Only record the first time we detect the song
        if song_name not in detected_songs:
            detected_songs[song_name] = start_ms / 1000  # Convert to seconds
    
    os.remove(temp_file)

# Sort results by the time each song first appeared
sorted_songs = sorted(detected_songs.items(), key=lambda x: x[1])
print("Detected songs in order of appearance:")
for song_name, start_time in sorted_songs:
    print(f"- {song_name} (starts at ~{round(start_time, 2)}s)")

Each method has its pros: silence splitting is cleanest for well-separated songs, modifying Dejavu gives you raw performance, and sliding windows works for messy files with no clear silence gaps. Pick the one that fits your use case!

内容的提问来源于stack exchange,提问作者Douglas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:12:38