如何利用Python Dejavu库检测单个音频文件中的多首歌曲?
Great question! I’ve dealt with this exact limitation of Dejavu before—out of the box, it’s optimized to return the first matching track it finds, which doesn’t work for multi-song files. Here are three practical solutions to modify its behavior and identify all tracks in your MP3:
1. Split the Audio on Silence & Recognize Chunks
The easiest approach is to split the MP3 into individual song chunks using silence detection, then run Dejavu on each chunk. This mimics how you’d manually split a mixed MP3, and plays nicely with Dejavu’s existing workflow.
Use a library like pydub to handle the audio splitting:
from pydub import AudioSegment from pydub.silence import split_on_silence import dejavu import os # Initialize Dejavu with your database config config = { "database": { "host": "localhost", "user": "root", "password": "", "database": "dejavu" } } djv = dejavu.Dejavu(config) # Load your multi-song MP3 audio = AudioSegment.from_mp3("mixed_songs.mp3") # Split audio into chunks based on silence (tune these values to your file) chunks = split_on_silence( audio, min_silence_len=600, # Minimum silence duration to trigger a split (ms) silence_thresh=-35, # Volume threshold for "silence" (dB) keep_silence=300 # Keep a small amount of silence around chunks ) # Recognize each chunk and collect unique results detected_songs = [] for idx, chunk in enumerate(chunks): temp_file = f"temp_chunk_{idx}.wav" chunk.export(temp_file, format="wav") match = djv.recognize("file", temp_file) if match and match["song_name"] not in [s["song_name"] for s in detected_songs]: detected_songs.append(match) # Clean up temporary files os.remove(temp_file) print("Detected songs in order:") for song in detected_songs: print(f"- {song['song_name']} (confidence: {song['confidence']})")
2. Modify Dejavu’s Core Recognition Logic
If you want to avoid splitting audio, you can tweak Dejavu’s source code to scan the entire file and collect all matches instead of stopping at the first one.
Navigate to the recognize.py file in your Dejavu installation (usually in dejavu/recognize.py) and modify the recognize_file method:
Original code (abbreviated):
# ... after calculating matches ... for song_id, offset in sorted(matches.items(), key=lambda x: x[1]): song = self.db.get_song_by_id(song_id) return { 'song_id': song_id, 'song_name': song['song_name'], 'confidence': len(matches[song_id]), 'offset': offset, 'offset_seconds': offset / self.FS, 'file_sha1': sha1 }
Modified code to return all matches:
# ... after calculating matches ... all_detected_songs = [] for song_id, offsets in matches.items(): song = self.db.get_song_by_id(song_id) # Calculate earliest offset for ordering earliest_offset = min(offsets) all_detected_songs.append({ 'song_id': song_id, 'song_name': song['song_name'], 'confidence': len(offsets), 'earliest_offset': earliest_offset, 'earliest_offset_seconds': earliest_offset / self.FS, 'file_sha1': sha1 }) # Sort songs by their earliest appearance in the file all_detected_songs.sort(key=lambda x: x['earliest_offset_seconds']) return all_detected_songs
Now when you call djv.recognize("file", "mixed_songs.mp3"), it will return a list of all matched songs ordered by their start time.
3. Sliding Window Scan
For more control (e.g., if silence detection isn’t reliable), use a sliding window approach: scan the MP3 in overlapping segments, run Dejavu on each segment, and deduplicate results.
import dejavu from pydub import AudioSegment import os config = { "database": { "host": "localhost", "user": "root", "password": "", "database": "dejavu" } } djv = dejavu.Dejavu(config) audio = AudioSegment.from_mp3("mixed_songs.mp3") window_length = 12 * 1000 # 12-second scan window step_length = 4 * 1000 # Move window forward by 4 seconds each time total_length = len(audio) # Track detected songs and their first occurrence time detected_songs = {} for start_ms in range(0, total_length - window_length, step_length): end_ms = start_ms + window_length window = audio[start_ms:end_ms] temp_file = "temp_window.wav" window.export(temp_file, format="wav") match = djv.recognize("file", temp_file) if match: song_name = match["song_name"] # Only record the first time we detect the song if song_name not in detected_songs: detected_songs[song_name] = start_ms / 1000 # Convert to seconds os.remove(temp_file) # Sort results by the time each song first appeared sorted_songs = sorted(detected_songs.items(), key=lambda x: x[1]) print("Detected songs in order of appearance:") for song_name, start_time in sorted_songs: print(f"- {song_name} (starts at ~{round(start_time, 2)}s)")
Each method has its pros: silence splitting is cleanest for well-separated songs, modifying Dejavu gives you raw performance, and sliding windows works for messy files with no clear silence gaps. Pick the one that fits your use case!
内容的提问来源于stack exchange,提问作者Douglas

