You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将AI生成的词级时间戳映射到歌词文本?

歌词可视化工具:原歌词与AI转录时间戳的可靠映射方案

我正在开发一款歌词可视化工具,核心功能是通过计算音节语音相似度为每个音节分配押韵组,同组音节会以同色高亮显示。要实现和歌曲播放同步的交互式可视化,必须获取词级甚至音节级的时间戳。

目前已完成的工作:

  • 实现了单词拆分为音节并分配押韵组的算法,结果存储在text.json中
  • 使用whisper-timestamped库生成了包含AI转录歌词及对应时间戳的timestamps.json

核心问题:需要将timestamps.json中的时间戳整合到text.json中,但AI无法精准识别每一个单词,必须找到可靠的方法完成两者的单词映射。


数据示例

text.json(原歌词结构)

{
    "text": "Hey What do you say here",
    "verses": [
        {
            "text": "Hey What do you say here",
            "verse_id": 0,
            "bars": [
                {
                    "bar": "Hey What do you say here",
                    "bar_id": 0,
                    "rhyme_group": null,
                    "words": [
                        {
                            "word": "Hey",
                            "word_id": 0,
                            "syllables": [
                                {
                                    "syllable": "Hey",
                                    "syllable_id": 0,
                                    "pronunciation": [
                                        "HH EY1"
                                    ],
                                    "rhyme_group": 1,
                                    "timestamp": null
                                }
                            ]
                        },
                        {
                            "word": "What",
                            "word_id": 1,
                            "syllables": [
                                {
                                    "syllable": "What",
                                    "syllable_id": 1,
                                    "pronunciation": [
                                        "W AH1 T"
                                    ],
                                    "rhyme_group": 0,
                                    "timestamp": null
                                }
                            ]
                        },
                        {
                            "word": "do",
                            "word_id": 2,
                            "syllables": [
                                {
                                    "syllable": "do",
                                    "syllable_id": 2,
                                    "pronunciation": [
                                        "D UW1"
                                    ],
                                    "rhyme_group": 2,
                                    "timestamp": null
                                }
                            ]
                        },
                        {
                            "word": "you",
                            "word_id": 3,
                            "syllables": [
                                {
                                    "syllable": "you",
                                    "syllable_id": 3,
                                    "pronunciation": [
                                        "Y UW1"
                                    ],
                                    "rhyme_group": 2,
                                    "timestamp": null
                                }
                            ]
                        },
                        {
                            "word": "say",
                            "word_id": 4,
                            "syllables": [
                                {
                                    "syllable": "say",
                                    "syllable_id": 4,
                                    "pronunciation": [
                                        "S EY1"
                                    ],
                                    "rhyme_group": 1,
                                    "timestamp": null
                                }
                            ]
                        },
                        {
                            "word": "here",
                            "word_id": 5,
                            "syllables": [
                                {
                                    "syllable": "here",
                                    "syllable_id": 5,
                                    "pronunciation": [
                                        "HH IY1 R"
                                    ],
                                    "rhyme_group": 0,
                                    "timestamp": null
                                }
                            ]
                        }
                    ]
                }
            ]
        }
    ]
}

timestamps.json(AI转录时间戳结构)

{
  "text": "Hey What do you say here",
  "segments": [
    {
      "id": 0,
      "seek": 0,
      "start": 0.5,
      "end": 1.2,
      "text": " Hey!",
      "tokens": [ 25431, 2298 ],
      "temperature": 0.0,
      "avg_logprob": -0.6674491882324218,
      "compression_ratio": 0.8181818181818182,
      "no_speech_prob": 0.10241222381591797,
      "confidence": 0.51,
      "words": [
        {
          "text": "Hey!",
          "start": 0.5,
          "end": 1.2,
          "confidence": 0.51
        }
      ]
    },
    {
      "id": 1,
      "seek": 200,
      "start": 2.02,
      "end": 4.48,
      "text": " What do you say here?",
      "tokens": [ 50364, 4410, 12, 384, 631, 2630, 18146, 3610, 2506, 50464 ],
      "temperature": 0.0,
      "avg_logprob": -0.43492694334550336,
      "compression_ratio": 0.7714285714285715,
      "no_speech_prob": 0.06502953916788101,
      "confidence": 0.595,
      "words": [
        {
          "text": "What",
          "start": 2.02,
          "end": 3.78,
          "confidence": 0.441
        },
        {
          "text": "do",
          "start": 3.78,
          "end": 3.84,
          "confidence": 0.948
        },
        {
          "text": "you",
          "start": 3.84,
          "end": 4.0,
          "confidence": 0.935
        },
        {
          "text": "ray",
          "start": 4.0,
          "end": 4.14,
          "confidence": 0.347
        },
        {
          "text": "here?",
          "start": 4.14,
          "end": 4.48,
          "confidence": 0.998
        }
      ]
    }
  ],
  "language": "en"
}

需处理的四类映射场景

  • 完全匹配:原文本与AI转录词完全一致,直接映射时间戳即可
  • 识别错误词:AI识别的词不正确,但词的数量和原文本一致,可直接按顺序对齐时间戳
  • 多词识别为一词:AI将原文本多个词合并识别为单个词,仅能获取该合并词的首尾时间戳,难以拆分映射到原多个词
  • 单词识别为多词:AI将原文本单个词拆分为多个词识别,需将这些拆分词的首尾时间戳合并为原词的时间戳

已尝试方法与痛点

试过用Levenshtein距离等模糊字符串匹配方法,但无法可靠解决上述场景2-4的冲突,长文本下问题更突出。求一套能覆盖所有场景的可靠映射方案。

内容的提问来源于stack exchange,提问作者paulpelikan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 02:12:09