Xamarin Android中如何为TextToSpeech与SpeechRecognizer集成AcousticEchoCanceler
解决方案
根因说明
当前使用的SpeechRecognizer默认采用通用语音识别音频配置,未启用回声消除(AEC)能力,会直接拾取设备自身扬声器播放的TTS内容,触发识别回调导致误停止TTS。
前置配置
首先在AndroidManifest.xml中添加必需权限:
<uses-permission android:name="android.permission.RECORD_AUDIO" /> <uses-permission android:name="android.permission.MODIFY_AUDIO_SETTINGS" />
步骤1:启用系统回声消除能力
方案A:优先使用通信音频模式(适配90%以上设备)
修改SpeechService的两处代码:
- 调整
CreateSpeechIntent方法,指定音频源为通话模式,系统会自动开启AEC、降噪等预处理:
protected virtual Intent CreateSpeechIntent(bool partialResults) { var intent = new Intent(RecognizerIntent.ActionRecognizeSpeech); intent.PutExtra(RecognizerIntent.ExtraLanguagePreference, Java.Util.Locale.Default); intent.PutExtra(RecognizerIntent.ExtraLanguage, Java.Util.Locale.Default); intent.PutExtra(RecognizerIntent.ExtraLanguageModel, RecognizerIntent.LanguageModelFreeForm); intent.PutExtra(RecognizerIntent.ExtraCallingPackage, Application.Context.PackageName); intent.PutExtra(RecognizerIntent.ExtraPartialResults, partialResults); // 新增:指定音频源为语音通信模式,自动启用AEC intent.PutExtra(RecognizerIntent.ExtraAudioSource, (int)Android.Media.AudioSource.VoiceCommunication); return intent; }
- 调整
Listen方法,启动识别前切换音频模式,释放时恢复默认模式,同时增加短结果过滤:
[Obsolete] public IObservable<string> Listen(Action action = null) => Observable.Create<string>(ob => { var audioManager = (AudioManager)Application.Context.GetSystemService(Context.AudioService); // 切换到通信模式触发系统AEC audioManager.Mode = Mode.InCommunication; speechRecognizer = SpeechRecognizer.CreateSpeechRecognizer(Application.Context); var listener = new SpeechRecognitionListener(); listener.ReadyForSpeech = () => this.ListenSubject.OnNext(true); listener.PartialResults = sentence => { lock (this.syncLock) { sentence = sentence.Trim(); // 新增:过滤空结果和过短的识别结果,减少误触发 if (string.IsNullOrEmpty(sentence) || sentence.Length < 2) return; if (action != null) { action.Invoke(); } if (currentIndex > sentence.Length) currentIndex = 0; var newPart = sentence.Substring(currentIndex); currentIndex = sentence.Length; final = sentence; } }; listener.EndOfSpeech = () => { ob.OnNext(final); ob.OnCompleted(); this.ListenSubject.OnNext(false); }; speechRecognizer.SetRecognitionListener(listener); speechRecognizer.StartListening(this.CreateSpeechIntent(true)); return () => { audioManager.SetStreamMute(Stream.Notification, false); // 恢复默认音频模式 audioManager.Mode = Mode.Normal; stop = true; speechRecognizer?.StopListening(); speechRecognizer?.Destroy(); this.ListenSubject.OnNext(false); }; });
方案B:显式启用AcousticEchoCanceler(适配不支持自动AEC的设备)
如果方案A在部分设备不生效,需要自定义音频采集链路,显式初始化AEC:
- 先检测设备AEC支持性:
private bool IsAecSupported => AcousticEchoCanceler.IsAvailable;
- 创建自定义
AudioRecord实例,绑定AEC:
private AudioRecord InitAudioRecordWithAec() { int sampleRate = 16000; var channelConfig = ChannelIn.Mono; var audioFormat = Encoding.Pcm16bit; int bufferSize = AudioRecord.GetMinBufferSize(sampleRate, channelConfig, audioFormat); var audioRecord = new AudioRecord(AudioSource.VoiceCommunication, sampleRate, channelConfig, audioFormat, bufferSize); if (IsAecSupported) { var aec = AcousticEchoCanceler.Create(audioRecord.AudioSessionId); aec.Enabled = true; // 同时可以启用噪声抑制进一步优化识别效果 if (NoiseSuppressor.IsAvailable) { var ns = NoiseSuppressor.Create(audioRecord.AudioSessionId); ns.Enabled = true; } } return audioRecord; }
后续将AudioRecord采集的PCM数据流喂给SpeechRecognizer即可,该方案需要自己实现音频帧的写入逻辑,兼容性更强。
步骤2:添加业务层双重校验
修改MyViewModel的触发逻辑,增加相似度校验,避免极端场景的误触发:
// 存储当前正在播报的TTS文本,用于相似度校验 private string _currentSpeakText; public MyViewModel () { var speakText = "Xamarin is a Microsoft-owned San Francisco-based software company founded in May 2011 by the engineers that created Mono, Xamarin.Android and Xamarin.iOS, which are cross-platform implementations of the Common Language Infrastructure and Common Language Specifications."; _currentSpeakText = speakText; speak(speakText); Action actionAfterSpeechDetect = delegate { if (textToSpeechCancellationToken != null && !textToSpeechCancellationToken.IsCancellationRequested) { // 新增:和当前播报的TTS文本做相似度匹配,避免回声误触发 var similarity = GetSimilarity(final, _currentSpeakText); if (similarity < 0.7) // 相似度低于70%才判定是用户输入 { textToSpeechCancellationToken.Cancel(); } } }; using (var cancelSrc = new CancellationTokenSource()) { output = await DependencyService.Get<ISpeechRecognizer>().Listen(actionAfterSpeechDetect).ToTask(cancelSrc.Token); } } // 简易文本相似度计算方法 private double GetSimilarity(string str1, string str2) { if (string.IsNullOrEmpty(str1) || string.IsNullOrEmpty(str2)) return 0; int sameCharCount = str1.Intersect(str2).Count(); return (double)sameCharCount / Math.Max(str1.Length, str2.Length); }
内容的提问来源于stack exchange,提问作者Thamotharan G
相关产品推荐
相关产品推荐

