You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SpeechRecognizer后SpeechSynthesizer无法发声的问题求助

问题:使用SFSpeechRecognizer后AVSpeechSynthesizer停止发声

在SwiftUI(Xcode 15.2)、iOS 17.3.1(iPhone 14 Pro)环境下,使用SpeechSynthesizer与SpeechRecognizer完成识别任务后,SpeechSynthesizer停止输出声音。此前相关旧帖的解决方案均不适用,未调用SpeechRecognizer的初始化方法时,SpeechSynthesizer可正常工作。

我的SpeechSynthesizer实现代码

func speak(_ text: String) {
    let utterance = AVSpeechUtterance(string: text)
    utterance.voice = AVSpeechSynthesisVoice(identifier: self.appState.chatParameters.voiceIdentifer)
    utterance.rate = 0.5
    speechSynthesizer.speak(utterance)
}

SpeechRecognizer初始化代码

private static func prepareEngine() throws -> (AVAudioEngine, SFSpeechAudioBufferRecognitionRequest) {
    print("prepareEngine()")
    let audioEngine = AVAudioEngine()
    
    let request = SFSpeechAudioBufferRecognitionRequest()
    request.shouldReportPartialResults = false
    request.requiresOnDeviceRecognition = true

    let audioSession = AVAudioSession.sharedInstance()
    try audioSession.setCategory(.playAndRecord)
    try audioSession.setActive(true, options: .notifyOthersOnDeactivation)
    let inputNode = audioEngine.inputNode
    
    let recordingFormat = inputNode.outputFormat(forBus: 0)
    inputNode.installTap(onBus: 0, bufferSize: 1024, format: recordingFormat) {
        (buffer: AVAudioPCMBuffer, when: AVAudioTime) in
        request.append(buffer)
    }
    audioEngine.prepare()
    try audioEngine.start()
    
    return (audioEngine, request)
}

解决方案

问题核心是语音识别结束后,音频会话配置未正确恢复,导致语音合成无法获取音频输出权限。针对iOS 17的特性,可按以下步骤修复:

1. 完成识别后清理语音识别资源

在语音识别任务结束时,必须停止音频引擎、移除输入节点监听,并关闭当前激活的音频会话:

func stopRecognition() {
    audioEngine.stop()
    audioEngine.inputNode.removeTap(onBus: 0)
    do {
        try AVAudioSession.sharedInstance().setActive(false, options: .notifyOthersOnDeactivation)
    } catch {
        print("关闭音频会话失败: \(error)")
    }
}

2. 语音合成前重新配置音频会话

修改speak方法,在合成前将音频会话切换为播放模式:

func speak(_ text: String) {
    do {
        let audioSession = AVAudioSession.sharedInstance()
        try audioSession.setCategory(.playback)
        try audioSession.setActive(true)
    } catch {
        print("配置语音合成音频会话失败: \(error)")
    }
    
    let utterance = AVSpeechUtterance(string: text)
    utterance.voice = AVSpeechSynthesisVoice(identifier: self.appState.chatParameters.voiceIdentifer)
    utterance.rate = 0.5
    speechSynthesizer.speak(utterance)
    
    // 可选:设置代理以便后续恢复会话(如需再次识别)
    speechSynthesizer.delegate = self
}

3. 实现代理处理合成后的会话恢复

如果后续需要再次进行语音识别,可在合成完成后关闭播放模式的音频会话:

extension YourClass: AVSpeechSynthesizerDelegate {
    func speechSynthesizer(_ synthesizer: AVSpeechSynthesizer, didFinish utterance: AVSpeechUtterance) {
        do {
            try AVAudioSession.sharedInstance().setActive(false)
        } catch {
            print("语音合成后关闭音频会话失败: \(error)")
        }
    }
}

4. iOS 17权限注意事项

iOS 17对音频会话的激活逻辑更严格,切换模式时确保没有其他音频进程占用资源;开启requiresOnDeviceRecognition时,需确保已获取麦克风权限,但该设置不会直接影响语音合成功能。

内容的提问来源于stack exchange,提问作者Scott T. Miller

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 14:22:45