如何将AI生成的词级时间戳映射到歌词文本?
歌词可视化工具:原歌词与AI转录时间戳的可靠映射方案
我正在开发一款歌词可视化工具,核心功能是通过计算音节语音相似度为每个音节分配押韵组,同组音节会以同色高亮显示。要实现和歌曲播放同步的交互式可视化,必须获取词级甚至音节级的时间戳。
目前已完成的工作:
- 实现了单词拆分为音节并分配押韵组的算法,结果存储在
text.json中 - 使用
whisper-timestamped库生成了包含AI转录歌词及对应时间戳的timestamps.json
核心问题:需要将timestamps.json中的时间戳整合到text.json中,但AI无法精准识别每一个单词,必须找到可靠的方法完成两者的单词映射。
数据示例
text.json(原歌词结构)
{ "text": "Hey What do you say here", "verses": [ { "text": "Hey What do you say here", "verse_id": 0, "bars": [ { "bar": "Hey What do you say here", "bar_id": 0, "rhyme_group": null, "words": [ { "word": "Hey", "word_id": 0, "syllables": [ { "syllable": "Hey", "syllable_id": 0, "pronunciation": [ "HH EY1" ], "rhyme_group": 1, "timestamp": null } ] }, { "word": "What", "word_id": 1, "syllables": [ { "syllable": "What", "syllable_id": 1, "pronunciation": [ "W AH1 T" ], "rhyme_group": 0, "timestamp": null } ] }, { "word": "do", "word_id": 2, "syllables": [ { "syllable": "do", "syllable_id": 2, "pronunciation": [ "D UW1" ], "rhyme_group": 2, "timestamp": null } ] }, { "word": "you", "word_id": 3, "syllables": [ { "syllable": "you", "syllable_id": 3, "pronunciation": [ "Y UW1" ], "rhyme_group": 2, "timestamp": null } ] }, { "word": "say", "word_id": 4, "syllables": [ { "syllable": "say", "syllable_id": 4, "pronunciation": [ "S EY1" ], "rhyme_group": 1, "timestamp": null } ] }, { "word": "here", "word_id": 5, "syllables": [ { "syllable": "here", "syllable_id": 5, "pronunciation": [ "HH IY1 R" ], "rhyme_group": 0, "timestamp": null } ] } ] } ] } ] }
timestamps.json(AI转录时间戳结构)
{ "text": "Hey What do you say here", "segments": [ { "id": 0, "seek": 0, "start": 0.5, "end": 1.2, "text": " Hey!", "tokens": [ 25431, 2298 ], "temperature": 0.0, "avg_logprob": -0.6674491882324218, "compression_ratio": 0.8181818181818182, "no_speech_prob": 0.10241222381591797, "confidence": 0.51, "words": [ { "text": "Hey!", "start": 0.5, "end": 1.2, "confidence": 0.51 } ] }, { "id": 1, "seek": 200, "start": 2.02, "end": 4.48, "text": " What do you say here?", "tokens": [ 50364, 4410, 12, 384, 631, 2630, 18146, 3610, 2506, 50464 ], "temperature": 0.0, "avg_logprob": -0.43492694334550336, "compression_ratio": 0.7714285714285715, "no_speech_prob": 0.06502953916788101, "confidence": 0.595, "words": [ { "text": "What", "start": 2.02, "end": 3.78, "confidence": 0.441 }, { "text": "do", "start": 3.78, "end": 3.84, "confidence": 0.948 }, { "text": "you", "start": 3.84, "end": 4.0, "confidence": 0.935 }, { "text": "ray", "start": 4.0, "end": 4.14, "confidence": 0.347 }, { "text": "here?", "start": 4.14, "end": 4.48, "confidence": 0.998 } ] } ], "language": "en" }
需处理的四类映射场景
- 完全匹配:原文本与AI转录词完全一致,直接映射时间戳即可
- 识别错误词:AI识别的词不正确,但词的数量和原文本一致,可直接按顺序对齐时间戳
- 多词识别为一词:AI将原文本多个词合并识别为单个词,仅能获取该合并词的首尾时间戳,难以拆分映射到原多个词
- 单词识别为多词:AI将原文本单个词拆分为多个词识别,需将这些拆分词的首尾时间戳合并为原词的时间戳
已尝试方法与痛点
试过用Levenshtein距离等模糊字符串匹配方法,但无法可靠解决上述场景2-4的冲突,长文本下问题更突出。求一套能覆盖所有场景的可靠映射方案。
内容的提问来源于stack exchange,提问作者paulpelikan
相关产品推荐
相关产品推荐

