You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python库Librosa实现语音跟踪以生成精准LRC歌词文件

解决歌词同步LRC生成的时间戳偏差问题

现有代码的核心问题在于误用音乐节拍检测替代语音活动检测:

  • librosa.beat.beat_track是针对音乐节拍(鼓点、节奏重音)设计的,完全不关注人声的起始时间,前奏阶段的节拍会被误判为歌词起始点。
  • np.linspace均匀分配时间戳的逻辑,忽略了歌词与人声的对应关系,导致时间戳完全不符合实际演唱节奏。

方案1:基于语音活动检测(VAD)的人声定位

核心思路

先检测音频中有人声的时间段,过滤前奏、间奏等无声音频段,再将歌词行与人声时间段的起始点一一对应,生成准确的时间戳。

代码实现(结合Librosa与WebRTC VAD)

先安装依赖:

pip install librosa webrtcvad numpy pydub

修改后的功能代码:

import librosa
import numpy as np
import webrtcvad
from pydub import AudioSegment

def generate_lrc_with_vad(audio_file, lyrics_file, output_lrc="song.lrc"):
    # 加载并清洗歌词(过滤空行)
    with open(lyrics_file, "r", encoding="utf-8") as f:
        lyrics = [line.strip() for line in f.readlines() if line.strip()]
    
    # 转换音频格式适配WebRTC VAD要求(16kHz单声道)
    audio = AudioSegment.from_file(audio_file)
    audio = audio.set_frame_rate(16000).set_channels(1)
    samples = np.array(audio.get_array_of_samples())
    sr = 16000

    # 初始化VAD(模式3为最严格检测,适合人声定位)
    vad = webrtcvad.Vad(3)
    frame_duration = 30  # ms,VAD仅支持10/20/30ms帧长
    frame_size = int(sr * frame_duration / 1000)

    # 逐帧检测人声
    is_speech = []
    for i in range(0, len(samples), frame_size):
        frame = samples[i:i+frame_size]
        if len(frame) < frame_size:
            break
        frame_bytes = frame.tobytes()
        is_speech.append(vad.is_speech(frame_bytes, sr))

    # 提取人声起始时间戳
    speech_times = []
    in_speech = False
    for idx, speech_flag in enumerate(is_speech):
        if speech_flag and not in_speech:
            start_time = idx * frame_duration / 1000
            speech_times.append(start_time)
            in_speech = True
        elif not speech_flag and in_speech:
            in_speech = False

    # 适配歌词与时间戳的数量匹配
    audio_duration = librosa.get_duration(filename=audio_file)
    if len(speech_times) < len(lyrics):
        # 若歌词数多于人声段,补全剩余时间戳到音频末尾
        remaining_times = np.linspace(
            speech_times[-1] if speech_times else 0,
            audio_duration,
            len(lyrics) - len(speech_times) + 1
        )[1:]
        speech_times.extend(remaining_times)
    elif len(speech_times) > len(lyrics):
        # 若人声段多于歌词,截断多余时间戳
        speech_times = speech_times[:len(lyrics)]

    # 生成LRC格式内容
    lrc_lines = []
    for time, line in zip(speech_times, lyrics):
        minutes = int(time // 60)
        seconds = int(time % 60)
        hundredths = int((time % 1) * 100)
        lrc_lines.append(f"[{minutes:02}:{seconds:02}.{hundredths:02}]{line}")

    # 写入LRC文件
    with open(output_lrc, "w", encoding="utf-8") as f:
        f.write("\n".join(lrc_lines))

方案2:使用专业音频-文本对齐工具(Aeneas)

Aeneas是专门用于音频与文本对齐的工具,内置成熟的语音识别与对齐逻辑,准确率远高于手动实现,能直接生成标准LRC文件。

安装依赖:

pip install aeneas

示例代码:

from aeneas.executetask import ExecuteTask
from aeneas.task import Task

def generate_lrc_with_aeneas(audio_file, lyrics_file, output_lrc="song.lrc"):
    # 配置任务参数:语言类型、文本格式、输出格式
    config_string = u"task_language=zho|is_text_type=plain|os_task_file_format=lrc"
    task = Task(config_string=config_string)
    task.audio_file_path = audio_file
    task.text_file_path = lyrics_file
    task.sync_map_file_path = output_lrc

    # 执行对齐任务
    ExecuteTask(task).execute()

提示:需根据歌曲语言修改task_language参数,如英文用eng,日文用jpn。


音频处理优化建议

  • 降噪预处理:使用noisereduce库对音频降噪,减少背景噪音对VAD检测的干扰。
  • 歌词格式规范:确保歌词文件每一行对应一句完整演唱内容,避免空行或多行合并,否则会导致对齐偏差。
  • 人声增强:通过FFmpeg等工具提升音频中人声的比例,进一步提高检测准确率。

内容的提问来源于stack exchange,提问作者Bajomo Richard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 15:39:23