如何让Android SpeechRecognizer忽略TTS输出,优先识别用户语音?
问题:Android中TTS播放时让SpeechRecognizer忽略TTS输出,仅识别用户语音
我使用compileSdk 34、minSdk 33开发,单独使用TextToSpeech(TTS)和SpeechRecognizer时功能正常,但当两者几乎同时启动时,SpeechRecognizer只会识别TTS的输出,完全忽略我在这段时间的语音输入。我需要让识别器专注捕捉用户语音,过滤掉TTS的播放内容。
另外我注意到Android 13及以上机型具备这种能力:比如TTS播放时说出"Hey Google",我的语音会被识别,TTS自动暂停,Google助手开始监听。想知道如何在自己的APP中实现类似效果。
实验代码
package com.example.speechandspeak; import android.Manifest; import android.content.Intent; import android.content.pm.PackageManager; import android.os.Bundle; import android.speech.RecognitionListener; import android.speech.RecognizerIntent; import android.speech.SpeechRecognizer; import android.speech.tts.TextToSpeech; import android.util.Log; import android.widget.Toast; import androidx.appcompat.app.AppCompatActivity; import androidx.core.app.ActivityCompat; import androidx.core.content.ContextCompat; import java.util.ArrayList; import java.util.Locale; public class MainActivity extends AppCompatActivity implements RecognitionListener { private TextToSpeech tts; private SpeechRecognizer speechRecognizer; @Override protected void onCreate(Bundle savedInstanceState) { super.onCreate(savedInstanceState); setContentView(R.layout.activity_main); // Initialize TextToSpeech tts = new TextToSpeech(this, new TextToSpeech.OnInitListener() { @Override public void onInit(int status) { if (status == TextToSpeech.SUCCESS) { int ttsLang = tts.setLanguage(Locale.US); if (ttsLang == TextToSpeech.LANG_MISSING_DATA || ttsLang == TextToSpeech.LANG_NOT_SUPPORTED) { Toast.makeText(MainActivity.this, "Language is not supported!", Toast.LENGTH_SHORT).show(); } else { speakText(); } } else { Toast.makeText(MainActivity.this, "TTS Initialization failed!", Toast.LENGTH_SHORT).show(); } } }); // Check if microphone permission is granted if (ContextCompat.checkSelfPermission(this, android.Manifest.permission.RECORD_AUDIO) != PackageManager.PERMISSION_GRANTED) { // Permission is not granted, request it ActivityCompat.requestPermissions(this, new String[]{Manifest.permission.RECORD_AUDIO}, 1); } // Initialize SpeechRecognizer speechRecognizer = SpeechRecognizer.createSpeechRecognizer(this); speechRecognizer.setRecognitionListener(this); } private void speakText() { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); String message = getString(R.string.tts_message); tts.speak(message, TextToSpeech.QUEUE_FLUSH, null, null); // After speaking, initiate speech recognition Intent speechRecognizerIntent = new Intent(RecognizerIntent.ACTION_RECOGNIZE_SPEECH); speechRecognizer.startListening(speechRecognizerIntent); } @Override public void onReadyForSpeech(Bundle params) { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); } @Override public void onBeginningOfSpeech() { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); } @Override public void onRmsChanged(float rmsdB) { //Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); } @Override public void onBufferReceived(byte[] bytes) { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); } @Override public void onEndOfSpeech() { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); } @Override public void onError(int i) { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName() + "(" + i + ")"); } @Override public void onEvent(int eventType, Bundle params) { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); } @Override public void onPartialResults(Bundle partialResults) { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); } @Override public void onResults(Bundle results) { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); ArrayList<String> matches = results.getStringArrayList(SpeechRecognizer.RESULTS_RECOGNITION); if (matches != null && !matches.isEmpty()) { Log.v(this.getLocalClassName(), "Matches: " + matches); String text = matches.get(0).toLowerCase(); // Convert to lowercase for case-insensitive comparison if (text.equals("stop")) { if (tts.isSpeaking()) { tts.stop(); speechRecognizer.stopListening(); // Stop listening for further commands Toast.makeText(MainActivity.this, "Stopped!", Toast.LENGTH_SHORT).show(); } } } } @Override protected void onDestroy() { super.onDestroy(); if (tts != null) { tts.shutdown(); } } }
Logcat调试信息
12:40:09.125 MainActivity D Entered: speakText 12:40:09.297 D Entered: onReadyForSpeech 12:40:09.761 D Entered: onBeginningOfSpeech 12:40:13.365 D Entered: onEndOfSpeech 12:40:13.394 D Entered: onResults 12:40:13.396 V Matches: [demonstration of speech recognition and text to speech say stop to quit] 12:40:14.524 ProfileInstaller D Installing profile for com.example.speechandspeak
解决方案
1. 配置TTS的AudioAttributes,标记语音合成类型
通过设置TTS的音频属性,让系统识别这是语音合成输出,从而在语音识别时自动应用回声消除:
// 在TTS初始化成功后添加以下配置 AudioAttributes audioAttributes = new AudioAttributes.Builder() .setUsage(AudioAttributes.USAGE_ASSISTANT) // 标记为助手类语音输出 .setContentType(AudioAttributes.CONTENT_TYPE_SPEECH) // 内容类型为语音 .build(); tts.setAudioAttributes(audioAttributes);
2. 给SpeechRecognizer启用回声消除
在启动识别的Intent中添加EXTRA_ENABLE_ECHO_CANCELLATION参数,强制启用回声消除功能:
private void speakText() { Log.d(this.getLocalClassName(), "Entered: " + Thread.currentThread().getStackTrace()[2].getMethodName()); String message = getString(R.string.tts_message); tts.speak(message, TextToSpeech.QUEUE_FLUSH, null, null); // 修改识别Intent,添加回声消除参数 Intent speechRecognizerIntent = new Intent(RecognizerIntent.ACTION_RECOGNIZE_SPEECH); speechRecognizerIntent.putExtra(RecognizerIntent.EXTRA_LANGUAGE_MODEL, RecognizerIntent.LANGUAGE_MODEL_FREE_FORM); speechRecognizerIntent.putExtra(RecognizerIntent.EXTRA_ENABLE_ECHO_CANCELLATION, true); // 启用回声消除 speechRecognizer.startListening(speechRecognizerIntent); }
3. 关于Google助手的实现原理
Google助手使用的是系统级的VoiceInteractionService,结合硬件级的回声消除技术,并且系统会持续监听唤醒词(如"Hey Google")。当检测到唤醒词时,系统会通过音频焦点抢占机制暂停正在播放的TTS,然后启动助手的识别流程。第三方APP无法直接使用系统级唤醒词监听,但通过上述两步配置,已经可以让SpeechRecognizer过滤掉自身TTS的输出,专注识别用户语音。
内容的提问来源于stack exchange,提问作者ususer
相关产品推荐
相关产品推荐

