使用AVSpeechSynthesizer做德语TTS时重复奇怪文本的问题
解决AVSpeechSynthesizer德语语音朗读时重复奇怪内容的问题
问题现象
使用AVSpeechSynthesizer实现文本转语音(TTS)功能,选择德语(de-DE)语音朗读以下文本时,会重复朗读„homograf“和„encore“这类无关内容:
ActualValue 3261,Day 30,Month July,Year 1990,Age 32,Star Sign Leo,Birthday 30 July 1990,Days since birth 11974,Days until birthday 79,Born on which day Monday,Birthday this year Sunday
用户实现代码如下:
var synth = AVSpeechSynthesizer() synth.delegate = self try? AVAudioSession.sharedInstance().setCategory(.playback, mode: .default, options: []) try? AVAudioSession.sharedInstance().setActive(true, options: .notifyOthersOnDeactivation) func playAudio(word: String){ let myUtterance = AVSpeechUtterance(string: word) myUtterance.voice = AVSpeechSynthesisVoice(language: "de-DE") if myUtterance.voice != nil { myUtterance.rate = Float(ttsConfig.speed)/10.0 myUtterance.preUtteranceDelay = 0.5 if isPauseAudio { synth.pauseSpeaking(at: .immediate) } else { synth.continueSpeaking() print("SPEAK EVENT----> \(word)") synth.speak(myUtterance) } } }
原因分析
- 德语语音引擎对混合语言文本(英文键名+数字)的解析出现异常,误将部分片段识别为需要特殊处理的同形异义词(homograf)或指令内容,触发重复朗读逻辑。
- 代码未针对德语语音做文本格式适配,直接传入混合语言文本,加剧了引擎的识别歧义。
解决方案
1. 文本预处理:替换为纯德语表述
将文本中的英文词汇全部替换为对应的德语词汇,消除语言混合带来的解析问题,示例转换后文本:
AktuellerWert 3261, Tag 30, Monat Juli, Jahr 1990, Alter 32, Sternzeichen Löwe, Geburtstag 30 Juli 1990, Tage seit Geburt 11974, Tage bis zum Geburtstag 79, Geburtswochentag Montag, Geburtstag in diesem Jahr Sonntag
2. 优化语音合成调用逻辑
确保每次朗读前终止未完成的任务,避免任务叠加导致的异常行为,修改后的playAudio函数:
func playAudio(word: String){ // 终止正在进行的朗读任务 if synth.isSpeaking { synth.stopSpeaking(at: .immediate) } let myUtterance = AVSpeechUtterance(string: word) guard let germanVoice = AVSpeechSynthesisVoice(language: "de-DE") else { print("德语语音不可用") return } myUtterance.voice = germanVoice myUtterance.rate = Float(ttsConfig.speed)/10.0 myUtterance.preUtteranceDelay = 0.5 if isPauseAudio { synth.pauseSpeaking(at: .immediate) } else { synth.continueSpeaking() print("SPEAK EVENT----> \(word)") synth.speak(myUtterance) } }
3. 避免混合语言文本直接传入
德语语音引擎对纯德语文本的解析稳定性更高,日常使用中尽量避免在德语语音模式下传入大量英文内容,减少识别歧义。
内容的提问来源于stack exchange,提问作者Tushar Premal
相关产品推荐
相关产品推荐

