You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Zynq-7000裸机运行TFLM时输出张量缓冲区不足致边界框异常

自定义TFLite模型在Zynq-7000裸机运行时的输出张量内存重叠问题

我将自行训练的自定义TFLite模型转换为C++源码,计划在Zynq-7000平台上以裸机方式运行。目前模型已成功运行,检测分数与Python端测试结果完全一致,但部分边界框数据存在异常。

板端运行结果

DEBUG: Detection 0 raw: score=0.8013, class=0.2714, box=[0.1784, 0.1757, 0.0000, 0.0000]
DEBUG: Detection 1 raw: score=0.5003, class=0.2192, box=[0.0000, 0.0000, 0.9921, 0.4017]
DEBUG: Detection 2 raw: score=0.3077, class=0.1971, box=[0.0039, 0.0133, 0.9910, 0.1075]
DEBUG: Detection 3 raw: score=0.2803, class=0.1810, box=[0.9604, 0.0426, 0.9986, 0.0785]
DEBUG: Detection 4 raw: score=0.2714, class=0.1784, box=[0.9536, 0.0270, 0.9939, 0.0553]
DEBUG: Detection 5 raw: score=0.2192, class=0.1757, box=[0.0408, 0.0082, 1.0023, 0.2196]
DEBUG: Detection 6 raw: score=0.1971, class=0.0000, box=[0.7465, 0.1062, 1.0030, 0.4046]
DEBUG: Detection 7 raw: score=0.1810, class=0.0000, box=[0.5294, 0.0006, 0.6696, 0.0322]
DEBUG: Detection 8 raw: score=0.1784, class=0.0000, box=[0.1593, 0.4119, 0.5908, 0.5678]
DEBUG: Detection 9 raw: score=0.1757, class=0.0000, box=[0.4822, 0.0018, 0.6665, 0.0422]

Python端测试结果

arr0_(scores) = [[0.8013255  0.5003249  0.30769435 0.28034142 0.27137667 0.219202
  0.19714184 0.1809863  0.17840956 0.17565072]]
arr1_(boxes) = [[
  [0.1093 0.2193 1.0176 0.7862]
  [0.7530 0.2164 0.9921 0.4017]
  [0.0039 0.0133 0.9910 0.1075]
  [0.9604 0.0426 0.9986 0.0785]
  [0.9536 0.0270 0.9939 0.0553]
  [0.0408 0.0082 1.0023 0.2196]
  [0.7465 0.1062 1.0030 0.4046]
  [0.5294 0.0006 0.6696 0.0322]
  [0.1593 0.4119 0.5908 0.5678]
  [0.4822 0.0018 0.6665 0.0422]
]]
arr2_num_detections = [10.]
arr3_class = [[0. 3. 0. 0. 0. 0. 0. 0. 0. 0.]]

可以看到,分数输出完全一致,但前两个边界框的前6个坐标存在差异,其余数据正常。

内存地址与张量分析

排查后认为问题源于输出张量的缓冲区不足,相关张量的内存地址如下:

tensor_score->data.f: 0x2654670
tensor_boxes->data.f: 0x2654690
tensor_count->data.f: 0x2654660
tensor_class->data.f: 0x2654680

通过Python端获取模型输出详情:

interpreter = tf.lite.Interpreter(model_path=TFLITE_MODEL_PATH)
interpreter.allocate_tensors()

output_details = interpreter.get_output_details()

for i, detail in enumerate(output_details):
    print(f"Output {i}: Name={detail['name']}, Shape={detail['shape']}, Type={detail['dtype']}")

输出结果:

Output 0: Name=StatefulPartitionedCall:1, Shape=[ 1 10], Type=<class 'numpy.float32'>
Output 1: Name=StatefulPartitionedCall:3, Shape=[ 1 10  4], Type=<class 'numpy.float32'>
Output 2: Name=StatefulPartitionedCall:0, Shape=[1], Type=<class 'numpy.float32'>
Output 3: Name=StatefulPartitionedCall:2, Shape=[ 1 10], Type=<class 'numpy.float32'>

这些输出张量的元素数量分别为10、40、1、10,类型为float32(每个元素占4字节),所需缓冲区大小至少为40、160、4、40字节。从板端的class输出可以看出内存重叠迹象:class的前6个值与score的后6个值一致,前两个边界框坐标与score的最后两个值一致。

有没有人在使用TFLM时遇到过类似问题?我曾尝试硬编码值测试,但仍无法找到解决方向。


内容的提问来源于stack exchange,提问作者caleb losch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 14:52:34