You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中实现强制对齐?Aeneas与P2FA报错求助

音频与转录文本对齐获取单词时间戳的问题求助

我需要将已有转录文本与音频对齐,获取每个单词的发音时间戳,但试用的两个模块均无法正常运行,不确定是操作有误还是环境配置问题,寻求更简便的实现方式,或解决以下报错:

1. Aeneas模块报错:aeneas.audiofile.AudioFileProbeError: Unable to call ffprobe executable

使用的代码:

audio_file = "audio.wav"
text_file = "text.txt"
syncmap_file = "syncmap.json"

config_string = f"task_language=en|os_task_file_format=json"
task = Task(config_string=config_string)
task.text_file_path_absolute = text_file
task.audio_file_path_absolute = audio_file
task.sync_map_file_path_absolute = syncmap_file

ExecuteTask(task).execute()
task.output_sync_map_file()

报错原因与解决方法

  • 原因:Aeneas依赖FFmpeg的ffprobe工具,系统未安装FFmpeg或ffprobe未加入系统环境变量PATH,导致程序无法调用。
  • 解决:
    1. 安装FFmpeg,确保ffprobe可执行文件所在目录被添加到系统PATH;
    2. 若不想修改环境变量,可在配置字符串中指定ffprobe的绝对路径,例如:
      config_string = f"task_language=en|os_task_file_format=json|ffprobe_path=/usr/bin/ffprobe"  # Windows下类似"C:\\ffmpeg\\bin\\ffprobe.exe"
      

2. P2FA模块报错:'HCopy' 未识别为内部或外部命令、'HVite' 未识别为内部或外部命令及文件找不到错误

使用的代码:

from p2fa_py3.p2fa import align

audio_path = "dstm.wav"
transcript_path = "text.txt"

output_dir = "syncmap.json"

align.align(audio_path, transcript_path, output_dir)

报错原因与解决方法

  • 原因:P2FA依赖HTK工具包(包含HCopy、HVite等工具),系统未安装HTK或工具未加入PATH;同时部分P2FA版本需要手动指定预训练模型文件路径,否则会出现文件找不到错误。
  • 解决:
    1. 下载安装HTK工具包,将HCopy、HVite所在目录添加到系统PATH;
    2. 检查P2FA的模型文件是否存在,若缺失需下载对应模型并在代码中指定模型路径(具体参考P2FA的文档)。

更简便的替代方案:使用OpenAI Whisper获取单词时间戳

Whisper是OpenAI开发的语音识别模型,可直接输出单词级时间戳,无需复杂的依赖配置,步骤如下:

  1. 安装依赖:
pip install openai-whisper
# 若系统未安装FFmpeg,需先安装(例如Ubuntu: sudo apt install ffmpeg; Windows: 下载FFmpeg并添加到PATH)
  1. 示例代码:
import whisper
import json

# 加载模型,可选择"base"、"small"、"medium"等,模型越大精度越高
model = whisper.load_model("base")
# 开启word_timestamps参数以获取单词时间戳
result = model.transcribe("audio.wav", word_timestamps=True)

# 提取并打印单词时间戳
for segment in result["segments"]:
    for word in segment["words"]:
        print(f"单词: {word['word'].strip()}, 开始: {word['start']:.2f}s, 结束: {word['end']:.2f}s")

# 保存为JSON文件
with open("word_timestamps.json", "w", encoding="utf-8") as f:
    # 仅保留包含单词时间戳的片段数据
    json.dump([{
        "start": seg["start"],
        "end": seg["end"],
        "words": [{
            "word": w["word"].strip(),
            "start": w["start"],
            "end": w["end"]
        } for w in seg["words"]]
    } for seg in result["segments"]], f, ensure_ascii=False, indent=2)

内容的提问来源于stack exchange,提问作者someoneontheglobe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 12:52:41