You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何借助Android Text-to-speech API实现流畅语音并跟踪当前朗读单词

Solution for Natural TTS Playback with Word-Level Tracking

Great question! I’ve run into this exact issue before—balancing natural, flowing text-to-speech with tracking the current spoken word is tricky, but there are two reliable approaches depending on your target Android API level.

1. For Android 11 (API 30) and Higher: Use SpeechTimingInfo

Android 11 introduced the getSpeechTimingInfo() method, which gives you precise timestamp data for every word in your text. This lets you track words in real-time without breaking the TTS engine’s natural intonation.

Step-by-Step Implementation:

val tts = TextToSpeech(context) { status ->
    if (status == TextToSpeech.SUCCESS) {
        tts.setLanguage(Locale.US)
        
        val targetText = "This is the sentence"
        val utteranceId = "tracking_utterance_1"
        val params = hashMapOf(
            TextToSpeech.Engine.KEY_PARAM_UTTERANCE_ID to utteranceId
        )

        // Get timing data before or right after starting playback
        tts.getSpeechTimingInfo(targetText, params)?.let { timingInfo ->
            // Store word timings to reference later
            val wordTimings = timingInfo.wordTimings
            val playbackStartTime = System.currentTimeMillis()

            // Use a handler to check current time and match to the active word
            val handler = Handler(Looper.getMainLooper())
            val checkWordRunnable = object : Runnable {
                override fun run() {
                    val elapsedTime = System.currentTimeMillis() - playbackStartTime
                    val currentWord = wordTimings.firstOrNull {
                        elapsedTime >= it.startTimeMs && elapsedTime <= it.endTimeMs
                    }?.word
                    
                    currentWord?.let {
                        // Update your UI or tracking logic here (e.g., highlight the word)
                        Log.d("TTS Tracking", "Current word: $it")
                    }

                    // Keep checking until playback ends
                    if (elapsedTime < timingInfo.totalAudioDurationMs) {
                        handler.postDelayed(this, 50) // Adjust interval for smoother tracking
                    }
                }
            }

            // Start playback and begin tracking
            tts.speak(targetText, TextToSpeech.QUEUE_FLUSH, params, utteranceId)
            handler.post(checkWordRunnable)
        }
    }
}

How it works:

  • getSpeechTimingInfo() returns a SpeechTimingInfo object with wordTimings—each entry includes the word text, start time, end time, and duration.
  • We use a Handler to periodically check the elapsed playback time and match it to the current word’s timestamp range.

2. For Pre-API 30: Use SSML Markers

If you need to support older devices, you can wrap each word in SSML <mark> tags. The TTS engine will trigger a callback when it reaches each marker, letting you track the current word without splitting the text into separate utterances.

Step-by-Step Implementation:

First, build SSML with word markers:

fun buildMarkedSsml(text: String): String {
    val words = text.split(" ")
    val ssmlBuilder = StringBuilder("<speak>")
    
    words.forEachIndexed { index, word ->
        // Add a marker before each word with a unique name
        ssmlBuilder.append("<mark name=\"word_$index\"/>$word ")
    }
    
    ssmlBuilder.append("</speak>")
    return ssmlBuilder.toString().trim()
}

Then play the SSML and track markers:

val tts = TextToSpeech(context) { status ->
    if (status == TextToSpeech.SUCCESS) {
        tts.setLanguage(Locale.US)
        val targetText = "This is the sentence"
        val markedSsml = buildMarkedSsml(targetText)
        val utteranceId = "ssml_tracking_utterance"
        
        val params = hashMapOf(
            TextToSpeech.Engine.KEY_PARAM_UTTERANCE_ID to utteranceId,
            TextToSpeech.Engine.KEY_PARAM_SSML to "true" // Tell TTS this is SSML
        )

        tts.setOnUtteranceProgressListener(object : UtteranceProgressListener() {
            override fun onStart(utteranceId: String?) {}
            override fun onDone(utteranceId: String?) {}
            override fun onError(utteranceId: String?) {}

            override fun onRangeStarted(utteranceId: String?, start: Int, end: Int, frame: Int) {
                // The frame parameter corresponds to the marker index
                val words = targetText.split(" ")
                val currentWord = words.getOrNull(frame)
                currentWord?.let {
                    // Update your tracking logic here
                    Log.d("TTS Tracking", "Current word: $it")
                }
            }
        })

        // Start natural playback of the full SSML text
        tts.speak(markedSsml, TextToSpeech.QUEUE_FLUSH, params, utteranceId)
    }
}

Why this works:

  • SSML lets you embed invisible markers in the text without disrupting the TTS engine’s natural flow (no awkward pauses between words).
  • The onRangeStarted callback triggers when the engine reaches each <mark> tag, and the frame parameter maps directly to the marker index we created.

Why Splitting Words Fails

When you split text into individual words and play them one by one, the TTS engine can’t apply natural prosody (like linking words together, adjusting tone, or adding appropriate pauses). This leads to that robotic, disjointed sound you noticed. Both solutions above keep the text as a single utterance, so the engine can handle all the natural language processing.

内容的提问来源于stack exchange,提问作者oflisback

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:01:54