You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Swift语音转文字仅逐句输出,如何实现转录内容追加而非覆盖历史文本?

Swift语音转文字仅逐句输出,如何实现转录内容追加而非覆盖历史文本?

我仔细看了你的代码,发现问题的核心是你误解了SFSpeechRecognitionResult的formattedString特性——这个属性返回的是从识别会话开始到当前的完整累积文本,不是每次新增的句子内容。所以你现在的逻辑会导致每次识别完成时,把整个历史文本再重复追加一遍,反而造成内容重复/覆盖的假象。咱们来调整几处关键逻辑就能解决问题:

现有代码的核心问题

  1. 当result.isFinal时,直接将newText(完整累积文本)追加到fullTranscriptionText,导致重复内容(比如第一次说"你好",full变成"你好. ";第二次说"世界",newText是"你好世界",full就变成"你好. 你好世界. ")
  2. 实时显示时,self.textView.text = self.fullTranscriptionText + currentSentenceText,而currentSentenceText是完整的newText,会和fullTranscriptionText重复拼接,导致显示混乱。

修正后的关键逻辑实现

首先,在你的类变量里新增一个属性,用来记录上一次最终识别完成时的转录段数(用来定位新增内容):

private var lastFinalSegmentCount = 0 // 记录上一次final时的转录段数量

然后修改recognitionTask里的处理逻辑,替换原来的代码:

recognitionTask = speechRecognizer?.recognitionTask(with: recognitionRequest) { result, error in
    if let result = result {
        let bestTranscription = result.bestTranscription
        let newText = bestTranscription.formattedString
        
        if newText != self.lastRecognizedText {
            if result.isFinal {
                // 提取这次final时新增的文本段
                let newSegments = bestTranscription.segments.dropFirst(self.lastFinalSegmentCount)
                let addedText = newSegments.map { $0.substring }.joined(separator: " ")
                
                // 追加新增文本到fullTranscriptionText
                if !addedText.isEmpty {
                    self.fullTranscriptionText += addedText + ". "
                    self.lastFinalSegmentCount = bestTranscription.segments.count // 更新段数记录
                }
                
                self.textView.text = self.fullTranscriptionText
                print("Final Transcription: \(self.fullTranscriptionText)")
            } else {
                // 实时显示:历史内容 + 当前实时新增的部分
                let currentAddedText = newText.replacingOccurrences(of: self.fullTranscriptionText, with: "")
                self.textView.text = self.fullTranscriptionText + currentAddedText
            }

            self.lastRecognizedText = newText
            
            // 自动滚动到最新内容
            let bottom = NSMakeRange(self.textView.text.count - 1, 1)
            self.textView.scrollRangeToVisible(bottom)
        }
    }

    if let error = error {
        print("Speech recognition error: \(error.localizedDescription)")
        self.handleError(error)
        self.stopListening()
    }
}

修正逻辑说明

  1. 新增内容精准提取:通过bestTranscription.segments获取每一段识别的独立内容,用lastFinalSegmentCount过滤掉已经追加到fullTranscriptionText的部分,只取新增的段拼接成addedText,彻底避免重复。
  2. 实时显示优化:实时显示时,把newText中已经包含在fullTranscriptionText的部分替换为空,只显示历史内容+当前正在说的实时内容,文本框内容连贯不重复。
  3. 状态同步更新:每次识别完成后更新lastFinalSegmentCount,确保下一次能正确定位新增内容。

你可以把这段修正后的代码替换原来的recognitionTask逻辑,然后测试连续说话的场景,应该就能实现内容的正常追加了。

备注:内容来源于stack exchange,提问作者Le'Anthony Howell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.15 08:44:58