You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SpeechSynthesisUtterance API时能否选中正在朗读的单词?

SpeechSynthesisUtterance API:实时高亮朗读单词的实现方案

嘿,刚好对这个API挺熟悉的,直接给你唠唠关键点:原生的SpeechSynthesisUtterance API并没有提供直接获取当前朗读单词或光标位置的事件,也没法自动选中正在朗读的内容。不过别担心,我们可以通过一些变通方法实现类似的高亮效果,下面给你详细拆解:

为什么原生做不到?

原生API只提供了几个基础生命周期事件:onstart(开始朗读)、onend(结束朗读)、onpause(暂停)、onresume(恢复),没有逐词朗读的回调钩子,所以没法直接拿到当前正在朗读的单词位置。

变通实现方法:基于时长估算的同步高亮

核心思路是把文本拆成单个单词,结合语速估算每个单词的朗读时长,用定时器同步触发单词高亮。具体步骤如下:

  1. 把要朗读的文本拆分成单词数组,方便逐个处理
  2. 结合你设置的rate(语速),估算每个单词的朗读时长(可按字符数或音节数调整)
  3. 用setTimeout在对应时间点触发单词高亮,同时把完整文本传给朗读实例

下面是基于你原有代码修改的可运行示例:

// 要朗读的目标文本
const targetText = 'Hello World this is a test of speech synthesis API';
// 拆分为单词数组
const wordList = targetText.split(' ');
// 初始化朗读实例
const msg = new SpeechSynthesisUtterance();
const voices = window.speechSynthesis.getVoices();
msg.voice = voices[10]; 
msg.voiceURI = 'native';
msg.volume = 1;
msg.rate = 1; // 语速范围0.1-10,数值越大读得越快
msg.pitch = 2;
msg.text = targetText;
msg.lang = 'en-US';

// 高亮指定下标的单词
function highlightWord(index) {
  // 清除之前的高亮状态
  document.querySelectorAll('.highlighted').forEach(el => el.classList.remove('highlighted'));
  // 高亮当前单词
  if (wordList[index]) {
    const wordElement = document.getElementById(`word-${index}`);
    wordElement?.classList.add('highlighted');
  }
}

// 将单词渲染到页面(方便后续高亮操作)
function renderWordsToPage() {
  const container = document.getElementById('text-container');
  wordList.forEach((word, idx) => {
    const span = document.createElement('span');
    span.id = `word-${idx}`;
    span.textContent = `${word} `;
    container.appendChild(span);
  });
}

// 估算每个单词的朗读时长(简单按字符数*基础时长,可根据实际语音调整)
const baseCharDuration = 150; // 单个字符的基础朗读时长(毫秒)
const wordDurations = wordList.map(word => {
  // 语速越快,时长越短,因此除以当前语速值
  return (word.length * baseCharDuration) / msg.rate;
});

// 朗读开始时启动高亮定时器
msg.onstart = function() {
  let currentWordIndex = 0;
  let accumulatedTime = 0;
  // 为每个单词设置触发高亮的定时器
  wordDurations.forEach(duration => {
    setTimeout(() => {
      highlightWord(currentWordIndex);
      currentWordIndex++;
    }, accumulatedTime);
    accumulatedTime += duration;
  });
};

msg.onend = function(e) {
  console.log('Finished in ' + e.elapsedTime + ' seconds.');
  // 朗读结束后清除所有高亮
  highlightWord(-1);
};

// 先把单词渲染到页面
renderWordsToPage();
// 启动朗读
speechSynthesis.speak(msg);

同时需要在页面中添加对应的CSS样式,让高亮效果更明显:

.highlighted {
  background-color: #ffeb3b;
  font-weight: 600;
}

#text-container {
  font-size: 24px;
  line-height: 2;
  margin: 20px;
}

注意事项

这种方法是基于时长估算的,不同语音(尤其是跨语言语音)的朗读节奏存在差异,所以估算的时长可能会有轻微偏差。如果需要绝对精准的逐词同步,可能得借助第三方TTS服务,这类服务通常会提供包含逐词时间戳的回调接口。

内容的提问来源于stack exchange,提问作者1.21 gigawatts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:11:35