使用SpeechSynthesisUtterance API时能否选中正在朗读的单词?
SpeechSynthesisUtterance API:实时高亮朗读单词的实现方案
嘿,刚好对这个API挺熟悉的,直接给你唠唠关键点:原生的SpeechSynthesisUtterance API并没有提供直接获取当前朗读单词或光标位置的事件,也没法自动选中正在朗读的内容。不过别担心,我们可以通过一些变通方法实现类似的高亮效果,下面给你详细拆解:
为什么原生做不到?
原生API只提供了几个基础生命周期事件:onstart(开始朗读)、onend(结束朗读)、onpause(暂停)、onresume(恢复),没有逐词朗读的回调钩子,所以没法直接拿到当前正在朗读的单词位置。
变通实现方法:基于时长估算的同步高亮
核心思路是把文本拆成单个单词,结合语速估算每个单词的朗读时长,用定时器同步触发单词高亮。具体步骤如下:
- 把要朗读的文本拆分成单词数组,方便逐个处理
- 结合你设置的
rate(语速),估算每个单词的朗读时长(可按字符数或音节数调整) - 用
setTimeout在对应时间点触发单词高亮,同时把完整文本传给朗读实例
下面是基于你原有代码修改的可运行示例:
// 要朗读的目标文本 const targetText = 'Hello World this is a test of speech synthesis API'; // 拆分为单词数组 const wordList = targetText.split(' '); // 初始化朗读实例 const msg = new SpeechSynthesisUtterance(); const voices = window.speechSynthesis.getVoices(); msg.voice = voices[10]; msg.voiceURI = 'native'; msg.volume = 1; msg.rate = 1; // 语速范围0.1-10,数值越大读得越快 msg.pitch = 2; msg.text = targetText; msg.lang = 'en-US'; // 高亮指定下标的单词 function highlightWord(index) { // 清除之前的高亮状态 document.querySelectorAll('.highlighted').forEach(el => el.classList.remove('highlighted')); // 高亮当前单词 if (wordList[index]) { const wordElement = document.getElementById(`word-${index}`); wordElement?.classList.add('highlighted'); } } // 将单词渲染到页面(方便后续高亮操作) function renderWordsToPage() { const container = document.getElementById('text-container'); wordList.forEach((word, idx) => { const span = document.createElement('span'); span.id = `word-${idx}`; span.textContent = `${word} `; container.appendChild(span); }); } // 估算每个单词的朗读时长(简单按字符数*基础时长,可根据实际语音调整) const baseCharDuration = 150; // 单个字符的基础朗读时长(毫秒) const wordDurations = wordList.map(word => { // 语速越快,时长越短,因此除以当前语速值 return (word.length * baseCharDuration) / msg.rate; }); // 朗读开始时启动高亮定时器 msg.onstart = function() { let currentWordIndex = 0; let accumulatedTime = 0; // 为每个单词设置触发高亮的定时器 wordDurations.forEach(duration => { setTimeout(() => { highlightWord(currentWordIndex); currentWordIndex++; }, accumulatedTime); accumulatedTime += duration; }); }; msg.onend = function(e) { console.log('Finished in ' + e.elapsedTime + ' seconds.'); // 朗读结束后清除所有高亮 highlightWord(-1); }; // 先把单词渲染到页面 renderWordsToPage(); // 启动朗读 speechSynthesis.speak(msg);
同时需要在页面中添加对应的CSS样式,让高亮效果更明显:
.highlighted { background-color: #ffeb3b; font-weight: 600; } #text-container { font-size: 24px; line-height: 2; margin: 20px; }
注意事项
这种方法是基于时长估算的,不同语音(尤其是跨语言语音)的朗读节奏存在差异,所以估算的时长可能会有轻微偏差。如果需要绝对精准的逐词同步,可能得借助第三方TTS服务,这类服务通常会提供包含逐词时间戳的回调接口。
内容的提问来源于stack exchange,提问作者1.21 gigawatts
相关产品推荐
相关产品推荐

