You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TFLite推理崩溃:张量内存莫名不可访问,如何避免重复分配?

解决TFLite推理时张量内存被回收的问题

核心原因分析

你遇到的问题根源是TFLite默认的Arena内存分配器,在AllocateTensors()执行后,若解释器对象或其内部内存管理模块被GC标记为可回收,部分中间张量的内存会被自动释放。输入张量因kTfLiteArenaRw的分配属性被保留,但中间张量被重置为kTfLiteMmapRo类型,说明原内存已被销毁并重新映射为只读权重区域。

具体解决方案

  • 确保Interpreter对象持有强引用
    检查代码中Interpreter实例的存储位置:必须将其放在全局变量、类成员变量等不会被GC轻易回收的地方,不能用函数内的局部临时变量。函数执行完毕后局部对象会被GC回收,连带销毁其管理的所有张量内存。

  • 禁用内存自动回收策略
    初始化解释器时,通过TfLiteInterpreterOptions配置关闭内存裁剪,强制保留所有张量的分配内存:

    TfLiteInterpreterOptions* options = TfLiteInterpreterOptionsCreate();
    // 禁止自动释放未使用的张量内存
    TfLiteInterpreterOptionsSetAllowBufferHandleOutput(options, false);
    TfLiteInterpreterOptionsSetNumThreads(options, 4); // 根据设备调整线程数
    Interpreter* interpreter = TfLiteInterpreterCreate(model, options);
    
  • 手动锁定中间张量内存
    对所有中间张量调用TfLiteTensorKeep()方法,阻止GC回收其内存:

    for (int i = 0; i < TfLiteInterpreterGetTensorCount(interpreter); ++i) {
      TfLiteTensor* tensor = TfLiteInterpreterGetTensor(interpreter, i);
      // 锁定非输入输出类的中间张量
      if (tensor->allocation_type != kTfLiteArenaRw && tensor->allocation_type != kTfLiteArenaRwPersistent) {
        TfLiteTensorKeep(tensor);
      }
    }
    

    若后续需要释放内存,可调用TfLiteTensorRelease(),持续推理场景下保持锁定即可。

  • 替换为持久化内存分配器
    使用TfLitePersistentBufferAllocator替换默认分配器,将张量内存分配在GC不会触及的持久化区域:

    TfLitePersistentBufferAllocator* allocator = TfLitePersistentBufferAllocatorCreate();
    TfLiteInterpreterOptionsSetCustomAllocator(options, allocator, TfLitePersistentBufferAllocatorAllocate, TfLitePersistentBufferAllocatorDeallocate);
    

验证步骤

  1. 在AllocateTensors()后和Invoke()前,再次检查张量的allocation_type与内存指针,确认无异常变更。
  2. 运行推理,验证是否仍存在内存访问崩溃问题。
  3. 监控内存占用,确保禁用自动回收后无内存泄漏(长期运行程序需定期清理废弃张量内存)。

内容的提问来源于stack exchange,提问作者Andrey Honich

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 23:45:01