You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Objective-C中SFSpeechRecognizer启用OnDeviceTranscript时转录结果不追加问题修复

问题修复方案

核心问题

你当前的代码每次回调都直接用新的formattedString覆盖textString,完全没处理增量转录结果。开启设备端识别(requiresOnDeviceRecognition = YES)时,中间结果的返回逻辑和云端有差异——云端会自动保留已转录内容,但设备端为了性能优化,可能只返回新增片段或重新生成文本,所以必须手动维护累计内容,不能直接替换。

修复步骤

  1. 维护一个变量保存历史转录结果和累计文本,用来对比每次回调的新结果
  2. 通过SFSpeechRecognitionTranscription的segments属性,找出新增的转录片段,只追加这部分内容,而不是替换整个字符串

修改后的代码示例

// 在你的类里定义成员变量,保存历史结果和累计文本
@interface YourClass ()
@property (nonatomic, strong) SFSpeechRecognitionResult *previousResult;
@property (nonatomic, strong) NSMutableString *accumulatedText;
@end

@implementation YourClass

- (void)startSpeechRecognition:(NSString *)filePath useDeviceMode:(BOOL)device {
    NSLocale *locale =[[NSLocale alloc] initWithLocaleIdentifier:[[NSLocale preferredLanguages] firstObject]];
    SFSpeechRecognizer* recognizer = [[SFSpeechRecognizer alloc] initWithLocale:locale];
    SFSpeechURLRecognitionRequest* request = [[SFSpeechURLRecognitionRequest alloc] initWithURL:[NSURL fileURLWithPath:filePath]];

    if(device){
        request.requiresOnDeviceRecognition = YES;
    }

    // 初始化累计文本和历史结果
    self.accumulatedText = [NSMutableString string];
    self.previousResult = nil;

    [recognizer recognitionTaskWithRequest:request resultHandler:^(SFSpeechRecognitionResult * _Nullable result, NSError * _Nullable error) {
        if (error) {
            convertedString(self.accumulatedText.copy);
            return;
        }

        SFSpeechRecognitionTranscription *currentTrans = result.bestTranscription;
        if (!self.previousResult) {
            // 第一次回调,直接把完整文本存进去
            [self.accumulatedText appendString:currentTrans.formattedString];
        } else {
            // 对比新旧结果的片段,只追加新增部分
            SFSpeechRecognitionTranscription *prevTrans = self.previousResult.bestTranscription;
            NSInteger oldCount = prevTrans.segments.count;
            for (NSInteger i = oldCount; i < currentTrans.segments.count; i++) {
                SFSpeechRecognitionSegment *segment = currentTrans.segments[i];
                [self.accumulatedText appendString:segment.substring];
                // 按需添加空格,保证文本连贯
                if (i < currentTrans.segments.count - 1) {
                    [self.accumulatedText appendString:@" "];
                }
            }
        }

        // 更新历史结果
        self.previousResult = result;

        // 实时返回累计的转录文本
        convertedString(self.accumulatedText.copy);

        if (result.isFinal) {
            // 最终结果返回
            convertedString(self.accumulatedText.copy);
        }
    }];
}

@end

关键说明

  • segments属性的作用:每个SFSpeechRecognitionSegment对应一段独立的转录内容,通过对比前后结果的片段数量,能精准拿到新增的文本,避免重复或覆盖
  • 设备端识别的特殊性:设备端识别优先考虑性能,不会每次都返回完整的已转录内容,必须手动维护累计字符串
  • UI更新注意:如果convertedString涉及UI操作,要把回调内容放到主线程执行,比如用dispatch_async(dispatch_get_main_queue(), ^{ ... })包裹

内容的提问来源于stack exchange,提问作者AJITHKUMAR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 04:12:56