如何利用语音识别结果实现元素筛选检索?纯JS朗读网站开发求助
Hey Florian, let’s break down your challenges step by step—no need to stress about being new to vanilla JS, we’ll tackle this one piece at a time!
Text into Elements
First, let’s solve the core task of wrapping each word in a <span>. This is straightforward with basic DOM manipulation:
// Target all <p> elements on the page document.querySelectorAll('p').forEach(paragraph => { // Get the text, trim extra spaces, and split into words (handles multiple spaces) const words = paragraph.textContent.trim().split(/\s+/); // Wrap each word in a <span> (add a class for styling/scripting later) const wrappedWords = words.map(word => `<span class="word">${word}</span>`).join(' '); // Replace the original paragraph content with the wrapped words paragraph.innerHTML = wrappedWords; });
trim()cleans up extra spaces at the start/end of the paragraphsplit(/\s+/)splits on one or more spaces, so you won’t get empty spans from multiple spaces- The
wordclass lets you target individual words for styling (like highlighting during speech) or logic later
To make the speech "wait" after specific words, we’ll use the Web Speech API’s SpeechSynthesisUtterance and its boundary events to track which word is being spoken:
const utterance = new SpeechSynthesisUtterance(); const wordSpans = document.querySelectorAll('.word'); // Set the utterance text to your full content (or a specific paragraph) utterance.text = document.body.textContent; // Track when the speech reaches a new word boundary utterance.onboundary = (event) => { // Clear active state from all words wordSpans.forEach(span => span.classList.remove('active')); // Find which span corresponds to the current speech position let currentCharIndex = 0; for (const span of wordSpans) { const spanLength = span.textContent.length; if (event.charIndex >= currentCharIndex && event.charIndex < currentCharIndex + spanLength) { span.classList.add('active'); // Mark the word being spoken // Add pause logic for specific words (e.g., "2") if (span.textContent === '2') { speechSynthesis.pause(); // Resume after 3 seconds (adjust as needed) setTimeout(() => speechSynthesis.resume(), 3000); } break; } currentCharIndex += spanLength + 1; // +1 accounts for the space between words } }; // Start speaking speechSynthesis.speak(utterance);
- The
onboundaryevent fires when the speech moves to a new word/syllable, so we can map that to our<span>elements - The pause logic checks if the current word matches your target (like "2") and pauses/resumes the speech synthesis
For voice control (like starting/stopping speech or searching), use the Web Speech API’s SpeechRecognition (note: it’s prefixed in Chrome as webkitSpeechRecognition):
// Initialize speech recognition const recognition = new (window.SpeechRecognition || window.webkitSpeechRecognition)(); recognition.continuous = true; // Keep listening after one command recognition.interimResults = false; // Only return final, confirmed results recognition.onresult = (event) => { const command = event.results[event.results.length - 1][0].transcript.trim().toLowerCase(); // Example commands switch(command) { case '开始朗读': speechSynthesis.resume(); break; case '暂停朗读': speechSynthesis.pause(); break; case '停止朗读': speechSynthesis.cancel(); break; default: // Handle search commands (e.g., "查找苹果") if (command.startsWith('查找')) { const searchTerm = command.replace('查找', '').trim(); // Highlight matching words wordSpans.forEach(span => { span.classList.toggle('highlight', span.textContent.toLowerCase().includes(searchTerm)); }); } } }; // Start listening recognition.start();
- This listens for voice commands and triggers actions like controlling speech or highlighting search terms
Since you already have a list, we can add a search input to filter items and link them to your text:
<!-- Add this to your HTML --> <input type="text" id="listSearch" placeholder="搜索章节..."> <ul id="contentList"> <li data-target="p1">第一章:引言</li> <li data-target="p2">第二章:主要内容</li> </ul>
const searchInput = document.getElementById('listSearch'); const listItems = document.querySelectorAll('#contentList li'); // Filter list items as user types searchInput.addEventListener('input', (e) => { const searchTerm = e.target.value.toLowerCase(); listItems.forEach(item => { item.style.display = item.textContent.toLowerCase().includes(searchTerm) ? 'block' : 'none'; }); }); // Click a list item to start reading the targeted paragraph listItems.forEach(item => { item.addEventListener('click', () => { const targetPara = document.getElementById(item.dataset.target); if (targetPara) { speechSynthesis.cancel(); // Stop current speech const newUtterance = new SpeechSynthesisUtterance(targetPara.textContent); speechSynthesis.speak(newUtterance); } }); });
- The
data-targetattribute links list items to specific<p>elements (give your paragraphsidattributes to match) - The input filters list items in real-time as the user types
Take it slow—start with the word wrapping first, then move to the speech sync logic, and tackle speech recognition and search once those are solid. You’ve already got the list done, which is a great start!
内容的提问来源于stack exchange,提问作者Florian Fromager

