You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Apple Watch音频流至iPhone用SFSpeechRecognizer无转录结果问题求助

实时语音识别跨设备传输无转录结果问题排查与解决思路

我需要在WatchOS应用中实现实时语音识别转录功能,由于WatchOS不支持SFSpeechRecognizer,因此通过WatchConnectivity将Apple Watch的音频流传输至iPhone伴侣应用。此前在iPhone单独测试相同的识别代码可正常工作,但通过流传输时,iPhone能接收音频块且无报错,却无法生成任何转录文本。怀疑是AVAudioPCMBuffer的转换环节出现问题,但因缺乏原始数据和指针处理经验,无法准确定位问题。以下是完整实现流程及对应代码:

实现流程及代码

  1. 用户点击按钮,触发Watch请求iPhone创建recognitionTask
// Watch端代码
func requestRecognitionTask() {
    guard WCSession.default.isReachable else { return }
    let message = ["action": "createRecognitionTask"]
    WCSession.default.sendMessage(message, replyHandler: { response in
        if response["success"] as? Bool == true {
            // 处理成功逻辑
        } else {
            // 处理错误
        }
    }, errorHandler: { error in
        print("请求创建识别任务失败:\(error)")
    })
}
  1. iPhone创建recognitionTask并返回成功或错误信息
// iPhone端代码
func session(_ session: WCSession, didReceiveMessage message: [String : Any], replyHandler: @escaping ([String : Any]) -> Void) {
    if message["action"] as? String == "createRecognitionTask" {
        do {
            let recognitionRequest = SFSpeechAudioBufferRecognitionRequest()
            recognitionRequest.shouldReportPartialResults = true
            recognitionTask = speechRecognizer.recognitionTask(with: recognitionRequest) { result, error in
                if let result = result {
                    print("转录结果:\(result.bestTranscription.formattedString)")
                }
                if error != nil || result?.isFinal == true {
                    recognitionTask?.finish()
                    recognitionTask = nil
                }
            }
            replyHandler(["success": true])
        } catch {
            replyHandler(["success": false, "error": error.localizedDescription])
        }
    }
}
  1. Watch配置音频会话,为音频引擎输入节点添加Tap并将音频格式发送至iPhone
// Watch端代码
func configureAudioSession() {
    do {
        try AVAudioSession.sharedInstance().setCategory(.playAndRecord, mode: .default)
        try AVAudioSession.sharedInstance().setActive(true)
        
        let audioEngine = AVAudioEngine()
        let inputNode = audioEngine.inputNode
        let inputFormat = inputNode.inputFormat(forBus: 0)
        
        // 发送音频格式到iPhone
        guard WCSession.default.isReachable else { return }
        let formatDict = [
            "sampleRate": inputFormat.sampleRate,
            "channelCount": inputFormat.channelCount,
            "bitDepth": inputFormat.bitDepth,
            "formatID": inputFormat.formatID.rawValue
        ] as [String : Any]
        WCSession.default.sendMessage(["action": "setAudioFormat", "format": formatDict], replyHandler: { response in
            if response["success"] as? Bool == true {
                // 添加Tap
                inputNode.installTap(onBus: 0, bufferSize: 1024, format: inputFormat) { buffer, time in
                    // 发送音频块到iPhone
                    self.sendAudioBuffer(buffer)
                }
            }
        }, errorHandler: { error in
            print("发送音频格式失败:\(error)")
        })
    } catch {
        print("配置音频会话失败:\(error)")
    }
}
  1. iPhone保存音频格式并返回成功或错误信息
// iPhone端代码
var audioFormat: AVAudioFormat?

func session(_ session: WCSession, didReceiveMessage message: [String : Any], replyHandler: @escaping ([String : Any]) -> Void) {
    if message["action"] as? String == "setAudioFormat" {
        guard let formatDict = message["format"] as? [String: Any],
              let sampleRate = formatDict["sampleRate"] as? Double,
              let channelCount = formatDict["channelCount"] as? Int,
              let bitDepth = formatDict["bitDepth"] as? Int,
              let formatIDRaw = formatDict["formatID"] as? UInt32 else {
            replyHandler(["success": false])
            return
        }
        let formatID = AudioFormatID(formatIDRaw)
        let streamDescription = AudioStreamBasicDescription(
            mSampleRate: sampleRate,
            mFormatID: formatID,
            mFormatFlags: AudioFormatFlags(bitDepth),
            mBytesPerPacket: 2 * channelCount,
            mFramesPerPacket: 1,
            mBytesPerFrame: 2 * channelCount,
            mChannelsPerFrame: UInt32(channelCount),
            mBitsPerChannel: UInt32(bitDepth),
            mReserved: 0
        )
        audioFormat = AVAudioFormat(streamDescription: &streamDescription)
        replyHandler(["success": true])
    }
}
  1. Watch启动音频引擎
// Watch端代码
func startAudioEngine() {
    do {
        try audioEngine.start()
    } catch {
        print("启动音频引擎失败:\(error)")
    }
}

// 发送音频块的方法
func sendAudioBuffer(_ buffer: AVAudioPCMBuffer) {
    guard WCSession.default.isReachable else { return }
    let data = Data(bytes: buffer.audioBufferList.pointee.mBuffers.mData!, count: Int(buffer.audioBufferList.pointee.mBuffers.mDataByteSize))
    WCSession.default.sendMessage(["action": "addAudioBuffer", "data": data], replyHandler: nil, errorHandler: { error in
        print("发送音频块失败:\(error)")
    })
}
  1. iPhone接收音频块并追加至recognitionRequest
// iPhone端代码
func session(_ session: WCSession, didReceiveMessage message: [String : Any], replyHandler: @escaping ([String : Any]) -> Void) {
    if message["action"] as? String == "addAudioBuffer" {
        guard let data = message["data"] as? Data,
              let format = audioFormat,
              let recognitionRequest = recognitionRequest else { return }
        
        let buffer = AVAudioPCMBuffer(pcmFormat: format, frameCapacity: AVAudioFrameCount(data.count) / format.streamDescription.pointee.mBytesPerFrame)
        buffer?.frameLength = buffer!.frameCapacity
        data.copyBytes(to: buffer!.audioBufferList.pointee.mBuffers.mData!, count: data.count)
        
        recognitionRequest.append(buffer!)
    }
}

可能的解决思路

1. 音频格式一致性校验

  • 强制Watch端使用SFSpeechRecognizer支持的格式,比如AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 16000, channels: 1, interleaved: false),避免默认格式不兼容。
  • iPhone端接收格式后,打印audioFormat的完整参数,与Watch端发送的参数逐一对比,确认转换后的格式无错误。

2. AVAudioPCMBuffer转换正确性检查

  • Watch端发送音频数据时,打印buffer.audioBufferList.pointee.mBuffers.mDataByteSize和发送的Data长度,确认数据完整未截断。
  • iPhone端重建AVAudioPCMBuffer时,确保frameCapacity计算准确:frameCapacity = AVAudioFrameCount(data.count) / UInt32(format.streamDescription.pointee.mBytesPerFrame),避免容量不足导致数据损坏。
  • 将iPhone端重建的AVAudioPCMBuffer保存为本地音频文件,播放验证是否为有效音频,排除传输或转换导致的数据损坏。

3. WatchConnectivity传输可靠性优化

  • 改用sendMessageData或transferFile传输音频块,这类API更适合流式数据,降低小数据包丢失概率。
  • 为每个音频块添加序列号,iPhone端按序列号顺序追加到识别请求,避免乱序导致识别逻辑失效。

4. 识别请求配置与错误日志

  • 确认recognitionRequest.shouldReportPartialResults设为true,确保实时返回部分转录结果。
  • 在recognitionTask的回调中补充error的详细日志,排查是否存在格式不支持、权限不足等隐藏错误。

5. 权限与会话状态检查

  • 确认iPhone端已申请并获取NSSpeechRecognitionUsageDescription和NSMicrophoneUsageDescription权限。
  • 监控WCSession的连接状态,确保传输过程中会话未断开,避免音频块丢失。

内容的提问来源于stack exchange,提问作者szagun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 05:32:50