移动端WebApp熄屏/后台时WebSpeech API语音合成持续运行方案咨询
Hey there! Let's break down your problem and figure out how to get your TTS WebApp working in the background like music apps do.
Can your WebApp achieve background playback like music apps?
Yes, but not with the raw window.speechSynthesis API alone. Here's why and how to fix it:
Why music apps work in the background
Music streaming apps rely on HTML5 <audio>/<video> elements combined with the Media Session API. Browsers grant special privileges to these media elements: as long as playback was initiated by a user gesture (like clicking a button), they can continue running even when the tab is in the background or the screen is off. This is because browsers treat media playback as a high-priority user task.
The problem with window.speechSynthesis
The Web Speech API's text-to-speech output is handled by the browser's built-in synthesizer, which doesn't use the standard media element pipeline. Browsers prioritize saving resources when a tab isn't active, so they pause non-media audio (like speech synthesis) to conserve battery and CPU. Your current NoSleep.js workaround keeps the screen on, but that's not ideal for user battery life.
Solution: Convert speech synthesis output to a playable audio blob
To get background playback, you need to route the speech synthesis output into an <audio> element. Here's a beginner-friendly step-by-step approach (note: requires modern browser support):
Capture the speech synthesis audio stream
Use the Web Audio API to capture the output ofspeechSynthesisand convert it into a Blob that the<audio>element can play:// Initialize Web Audio context (works across browsers) const audioContext = new (window.AudioContext || window.webkitAudioContext)(); const destination = audioContext.createMediaStreamDestination(); const mediaRecorder = new MediaRecorder(destination.stream); const audioChunks = []; // Collect audio data as it's generated mediaRecorder.ondataavailable = (event) => { if (event.data.size > 0) audioChunks.push(event.data); }; // When recording stops, create an audio element and play it mediaRecorder.onstop = () => { const audioBlob = new Blob(audioChunks, { type: 'audio/wav' }); const audioUrl = URL.createObjectURL(audioBlob); const audioElement = new Audio(audioUrl); // Play the audio (must be triggered by user action!) audioElement.play(); // Optional: Add system-level playback controls if ('mediaSession' in navigator) { navigator.mediaSession.metadata = new MediaMetadata({ title: 'TTS Playback', artist: 'Web Speech API', }); } }; // Start recording before speaking mediaRecorder.start(); // Speak your text const synth = window.speechSynthesis; const utterance = new SpeechSynthesisUtterance('Your long text content here'); // Stop recording once speech finishes utterance.onend = () => mediaRecorder.stop(); synth.speak(utterance);Note: If audio capture doesn't work on some browsers, try using a lightweight library like
responsivevoicethat can export TTS directly as audio, or test with Chrome (it has the most robust Web Speech API support).Ensure user-initiated playback
Browsers block automatic background audio for privacy and battery reasons, so make sure all playback starts with a user action (like clicking a "Start Reading" button). This is a non-negotiable policy.Use Media Session API (optional but recommended)
The Media Session API lets you add play/pause controls to the system's notification center or lock screen. This makes the background playback experience feel polished, just like music apps.
Does mounting speechSynthesis to a non-window object help?
Nope. window.speechSynthesis is a global browser API—assigning it to another variable or object doesn't change how the browser handles its audio output. The core issue is the type of audio being played, not where the API reference is stored.
Final tips for a beginner
- Test on multiple mobile browsers (Chrome, Safari) since support for Web Speech and Web Audio APIs can vary.
- Start with short text chunks to debug, then scale to longer content.
- Avoid relying on NoSleep.js long-term—it's a hack that drains battery. The audio element approach is the proper, user-friendly solution.
内容的提问来源于stack exchange,提问作者Yaksha

