You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Swift语音识别:暂停识别以便将结果朗读给用户

解决语音识别与AVSpeechSynthesizer的冲突问题

核心思路

在触发语音合成朗读前暂停语音识别会话,待朗读完成后重新启动识别。同时通过检测特定短语来触发这套暂停-恢复逻辑。

具体实现步骤

  1. 管理语音识别会话状态
    如果你用的是SFSpeechRecognizer,需维护SFSpeechAudioBufferRecognitionRequest和SFSpeechRecognitionTask实例。触发朗读前,取消或暂停当前识别任务:

    // 暂停/取消当前识别任务
    recognitionTask?.cancel()
    recognitionTask = nil
    
  2. 监听朗读完成事件
    让你的类遵循AVSpeechSynthesizerDelegate协议,实现朗读结束后的回调,重启语音识别:

    func speechSynthesizer(_ synthesizer: AVSpeechSynthesizer, didFinish utterance: AVSpeechUtterance) {
        // 调用自定义方法重启语音识别会话
        startSpeechRecognition()
    }
    
  3. 检测特定短语触发暂停
    在语音识别结果回调中,判断文本是否包含目标短语,若包含则触发朗读并暂停识别:

    func speechRecognitionTask(_ task: SFSpeechRecognitionTask, didRecognitionResult result: SFSpeechRecognitionResult) {
        guard let bestTranscription = result.bestTranscription.formattedString else { return }
        
        // 示例:识别到"确认内容"时触发验证朗读
        if bestTranscription.contains("确认内容") {
            // 根据业务逻辑提取需要朗读的验证文本
            let contentToVerify = extractVerificationContent(from: bestTranscription)
            // 暂停识别
            recognitionTask?.cancel()
            recognitionTask = nil
            // 启动朗读
            let utterance = AVSpeechUtterance(string: contentToVerify)
            speechSynthesizer.speak(utterance)
        }
    }
    
  4. 额外优化建议
    优先使用暂停识别任务的方式,而非直接关闭麦克风,这样能保证朗读结束后快速恢复识别流程,减少用户等待时间。

内容的提问来源于stack exchange,提问作者Brian Kalski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 10:04:53