Apple Watch音频流至iPhone用SFSpeechRecognizer无转录结果问题求助
实时语音识别跨设备传输无转录结果问题排查与解决思路
我需要在WatchOS应用中实现实时语音识别转录功能,由于WatchOS不支持SFSpeechRecognizer,因此通过WatchConnectivity将Apple Watch的音频流传输至iPhone伴侣应用。此前在iPhone单独测试相同的识别代码可正常工作,但通过流传输时,iPhone能接收音频块且无报错,却无法生成任何转录文本。怀疑是AVAudioPCMBuffer的转换环节出现问题,但因缺乏原始数据和指针处理经验,无法准确定位问题。以下是完整实现流程及对应代码:
实现流程及代码
- 用户点击按钮,触发Watch请求iPhone创建recognitionTask
// Watch端代码 func requestRecognitionTask() { guard WCSession.default.isReachable else { return } let message = ["action": "createRecognitionTask"] WCSession.default.sendMessage(message, replyHandler: { response in if response["success"] as? Bool == true { // 处理成功逻辑 } else { // 处理错误 } }, errorHandler: { error in print("请求创建识别任务失败:\(error)") }) }
- iPhone创建recognitionTask并返回成功或错误信息
// iPhone端代码 func session(_ session: WCSession, didReceiveMessage message: [String : Any], replyHandler: @escaping ([String : Any]) -> Void) { if message["action"] as? String == "createRecognitionTask" { do { let recognitionRequest = SFSpeechAudioBufferRecognitionRequest() recognitionRequest.shouldReportPartialResults = true recognitionTask = speechRecognizer.recognitionTask(with: recognitionRequest) { result, error in if let result = result { print("转录结果:\(result.bestTranscription.formattedString)") } if error != nil || result?.isFinal == true { recognitionTask?.finish() recognitionTask = nil } } replyHandler(["success": true]) } catch { replyHandler(["success": false, "error": error.localizedDescription]) } } }
- Watch配置音频会话,为音频引擎输入节点添加Tap并将音频格式发送至iPhone
// Watch端代码 func configureAudioSession() { do { try AVAudioSession.sharedInstance().setCategory(.playAndRecord, mode: .default) try AVAudioSession.sharedInstance().setActive(true) let audioEngine = AVAudioEngine() let inputNode = audioEngine.inputNode let inputFormat = inputNode.inputFormat(forBus: 0) // 发送音频格式到iPhone guard WCSession.default.isReachable else { return } let formatDict = [ "sampleRate": inputFormat.sampleRate, "channelCount": inputFormat.channelCount, "bitDepth": inputFormat.bitDepth, "formatID": inputFormat.formatID.rawValue ] as [String : Any] WCSession.default.sendMessage(["action": "setAudioFormat", "format": formatDict], replyHandler: { response in if response["success"] as? Bool == true { // 添加Tap inputNode.installTap(onBus: 0, bufferSize: 1024, format: inputFormat) { buffer, time in // 发送音频块到iPhone self.sendAudioBuffer(buffer) } } }, errorHandler: { error in print("发送音频格式失败:\(error)") }) } catch { print("配置音频会话失败:\(error)") } }
- iPhone保存音频格式并返回成功或错误信息
// iPhone端代码 var audioFormat: AVAudioFormat? func session(_ session: WCSession, didReceiveMessage message: [String : Any], replyHandler: @escaping ([String : Any]) -> Void) { if message["action"] as? String == "setAudioFormat" { guard let formatDict = message["format"] as? [String: Any], let sampleRate = formatDict["sampleRate"] as? Double, let channelCount = formatDict["channelCount"] as? Int, let bitDepth = formatDict["bitDepth"] as? Int, let formatIDRaw = formatDict["formatID"] as? UInt32 else { replyHandler(["success": false]) return } let formatID = AudioFormatID(formatIDRaw) let streamDescription = AudioStreamBasicDescription( mSampleRate: sampleRate, mFormatID: formatID, mFormatFlags: AudioFormatFlags(bitDepth), mBytesPerPacket: 2 * channelCount, mFramesPerPacket: 1, mBytesPerFrame: 2 * channelCount, mChannelsPerFrame: UInt32(channelCount), mBitsPerChannel: UInt32(bitDepth), mReserved: 0 ) audioFormat = AVAudioFormat(streamDescription: &streamDescription) replyHandler(["success": true]) } }
- Watch启动音频引擎
// Watch端代码 func startAudioEngine() { do { try audioEngine.start() } catch { print("启动音频引擎失败:\(error)") } } // 发送音频块的方法 func sendAudioBuffer(_ buffer: AVAudioPCMBuffer) { guard WCSession.default.isReachable else { return } let data = Data(bytes: buffer.audioBufferList.pointee.mBuffers.mData!, count: Int(buffer.audioBufferList.pointee.mBuffers.mDataByteSize)) WCSession.default.sendMessage(["action": "addAudioBuffer", "data": data], replyHandler: nil, errorHandler: { error in print("发送音频块失败:\(error)") }) }
- iPhone接收音频块并追加至recognitionRequest
// iPhone端代码 func session(_ session: WCSession, didReceiveMessage message: [String : Any], replyHandler: @escaping ([String : Any]) -> Void) { if message["action"] as? String == "addAudioBuffer" { guard let data = message["data"] as? Data, let format = audioFormat, let recognitionRequest = recognitionRequest else { return } let buffer = AVAudioPCMBuffer(pcmFormat: format, frameCapacity: AVAudioFrameCount(data.count) / format.streamDescription.pointee.mBytesPerFrame) buffer?.frameLength = buffer!.frameCapacity data.copyBytes(to: buffer!.audioBufferList.pointee.mBuffers.mData!, count: data.count) recognitionRequest.append(buffer!) } }
可能的解决思路
1. 音频格式一致性校验
- 强制Watch端使用
SFSpeechRecognizer支持的格式,比如AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 16000, channels: 1, interleaved: false),避免默认格式不兼容。 - iPhone端接收格式后,打印
audioFormat的完整参数,与Watch端发送的参数逐一对比,确认转换后的格式无错误。
2. AVAudioPCMBuffer转换正确性检查
- Watch端发送音频数据时,打印
buffer.audioBufferList.pointee.mBuffers.mDataByteSize和发送的Data长度,确认数据完整未截断。 - iPhone端重建
AVAudioPCMBuffer时,确保frameCapacity计算准确:frameCapacity = AVAudioFrameCount(data.count) / UInt32(format.streamDescription.pointee.mBytesPerFrame),避免容量不足导致数据损坏。 - 将iPhone端重建的
AVAudioPCMBuffer保存为本地音频文件,播放验证是否为有效音频,排除传输或转换导致的数据损坏。
3. WatchConnectivity传输可靠性优化
- 改用
sendMessageData或transferFile传输音频块,这类API更适合流式数据,降低小数据包丢失概率。 - 为每个音频块添加序列号,iPhone端按序列号顺序追加到识别请求,避免乱序导致识别逻辑失效。
4. 识别请求配置与错误日志
- 确认
recognitionRequest.shouldReportPartialResults设为true,确保实时返回部分转录结果。 - 在
recognitionTask的回调中补充error的详细日志,排查是否存在格式不支持、权限不足等隐藏错误。
5. 权限与会话状态检查
- 确认iPhone端已申请并获取
NSSpeechRecognitionUsageDescription和NSMicrophoneUsageDescription权限。 - 监控WCSession的连接状态,确保传输过程中会话未断开,避免音频块丢失。
内容的提问来源于stack exchange,提问作者szagun
相关产品推荐
相关产品推荐

