如何借助Android Text-to-speech API实现流畅语音并跟踪当前朗读单词
Great question! I’ve run into this exact issue before—balancing natural, flowing text-to-speech with tracking the current spoken word is tricky, but there are two reliable approaches depending on your target Android API level.
1. For Android 11 (API 30) and Higher: Use SpeechTimingInfo
Android 11 introduced the getSpeechTimingInfo() method, which gives you precise timestamp data for every word in your text. This lets you track words in real-time without breaking the TTS engine’s natural intonation.
Step-by-Step Implementation:
val tts = TextToSpeech(context) { status -> if (status == TextToSpeech.SUCCESS) { tts.setLanguage(Locale.US) val targetText = "This is the sentence" val utteranceId = "tracking_utterance_1" val params = hashMapOf( TextToSpeech.Engine.KEY_PARAM_UTTERANCE_ID to utteranceId ) // Get timing data before or right after starting playback tts.getSpeechTimingInfo(targetText, params)?.let { timingInfo -> // Store word timings to reference later val wordTimings = timingInfo.wordTimings val playbackStartTime = System.currentTimeMillis() // Use a handler to check current time and match to the active word val handler = Handler(Looper.getMainLooper()) val checkWordRunnable = object : Runnable { override fun run() { val elapsedTime = System.currentTimeMillis() - playbackStartTime val currentWord = wordTimings.firstOrNull { elapsedTime >= it.startTimeMs && elapsedTime <= it.endTimeMs }?.word currentWord?.let { // Update your UI or tracking logic here (e.g., highlight the word) Log.d("TTS Tracking", "Current word: $it") } // Keep checking until playback ends if (elapsedTime < timingInfo.totalAudioDurationMs) { handler.postDelayed(this, 50) // Adjust interval for smoother tracking } } } // Start playback and begin tracking tts.speak(targetText, TextToSpeech.QUEUE_FLUSH, params, utteranceId) handler.post(checkWordRunnable) } } }
How it works:
getSpeechTimingInfo()returns aSpeechTimingInfoobject withwordTimings—each entry includes the word text, start time, end time, and duration.- We use a
Handlerto periodically check the elapsed playback time and match it to the current word’s timestamp range.
2. For Pre-API 30: Use SSML Markers
If you need to support older devices, you can wrap each word in SSML <mark> tags. The TTS engine will trigger a callback when it reaches each marker, letting you track the current word without splitting the text into separate utterances.
Step-by-Step Implementation:
First, build SSML with word markers:
fun buildMarkedSsml(text: String): String { val words = text.split(" ") val ssmlBuilder = StringBuilder("<speak>") words.forEachIndexed { index, word -> // Add a marker before each word with a unique name ssmlBuilder.append("<mark name=\"word_$index\"/>$word ") } ssmlBuilder.append("</speak>") return ssmlBuilder.toString().trim() }
Then play the SSML and track markers:
val tts = TextToSpeech(context) { status -> if (status == TextToSpeech.SUCCESS) { tts.setLanguage(Locale.US) val targetText = "This is the sentence" val markedSsml = buildMarkedSsml(targetText) val utteranceId = "ssml_tracking_utterance" val params = hashMapOf( TextToSpeech.Engine.KEY_PARAM_UTTERANCE_ID to utteranceId, TextToSpeech.Engine.KEY_PARAM_SSML to "true" // Tell TTS this is SSML ) tts.setOnUtteranceProgressListener(object : UtteranceProgressListener() { override fun onStart(utteranceId: String?) {} override fun onDone(utteranceId: String?) {} override fun onError(utteranceId: String?) {} override fun onRangeStarted(utteranceId: String?, start: Int, end: Int, frame: Int) { // The frame parameter corresponds to the marker index val words = targetText.split(" ") val currentWord = words.getOrNull(frame) currentWord?.let { // Update your tracking logic here Log.d("TTS Tracking", "Current word: $it") } } }) // Start natural playback of the full SSML text tts.speak(markedSsml, TextToSpeech.QUEUE_FLUSH, params, utteranceId) } }
Why this works:
- SSML lets you embed invisible markers in the text without disrupting the TTS engine’s natural flow (no awkward pauses between words).
- The
onRangeStartedcallback triggers when the engine reaches each<mark>tag, and theframeparameter maps directly to the marker index we created.
Why Splitting Words Fails
When you split text into individual words and play them one by one, the TTS engine can’t apply natural prosody (like linking words together, adjusting tone, or adding appropriate pauses). This leads to that robotic, disjointed sound you noticed. Both solutions above keep the text as a single utterance, so the engine can handle all the natural language processing.
内容的提问来源于stack exchange,提问作者oflisback

