TensorFlow Lite转换后StatefulPartitionedCall:n输出含义解析及SSD-MobileNet迁移学习模型适配树莓派4的输出处理问题
Hey there! Let's break down your two main issues step by step—figuring out what those StatefulPartitionedCall:n outputs mean, and fixing that frustrating ValueError in your webcam detection code.
1. 映射TFLite输出到原模型的含义
When you convert a TensorFlow Object Detection API model to TFLite, the output names get renamed to generic StatefulPartitionedCall:n labels, but their structure and order align with the original SavedModel's outputs. For the ssd_mobilenet_v2_fpnlite_640x640_coco17_tpu-8 model, here's how to match each TFLite output to the original documented outputs, using the shape info you provided:
First, recall the original model's 8 standard outputs (from the Object Detection API):
num_detections: Float32 tensor of shape[batch_size]→ count of valid detectionsdetection_boxes: Float32 tensor of shape[batch_size, max_detections, 4]→ bounding box coordinates (ymin, xmin, ymax, xmax)detection_scores: Float32 tensor of shape[batch_size, max_detections]→ confidence scores for each detectiondetection_classes: Float32 tensor of shape[batch_size, max_detections]→ class IDs for each detectionraw_detection_boxes: Float32 tensor of shape[batch_size, num_anchors, 4]→ unfiltered boxes from all 51150 anchors (specific to this model)raw_detection_scores: Float32 tensor of shape[batch_size, num_anchors, num_classes]→ unfiltered scores from all anchorsdetection_multiclass_scores: Float32 tensor of shape[batch_size, max_detections, num_classes+1]→ multi-class scores for top detectionsdetection_anchor_indices: Float32 tensor of shape[batch_size, max_detections]→ anchor indices linked to top detections
Quick way to match your TFLite outputs:
For deployment, you only need the first four core outputs. Use shape to identify them:
- Shape
[1,1]→ This isnum_detections(your entry 1:StatefulPartitionedCall:4) - Shape
[1, X, 4](X is usually 100 for COCO models) →detection_boxes - Shape
[1, X]→ This will be eitherdetection_scoresordetection_classes - The other outputs are auxiliary and can be ignored for inference.
For absolute certainty, run the same test image through both your original .pb model and TFLite model, then compare output values (within floating-point tolerance). This will tell you exactly which StatefulPartitionedCall:n matches each original output.
2. Fixing the ValueError: The truth value of an array with more than one element is ambiguous
This error happens because your code is trying to evaluate a boolean condition on an entire array instead of individual scalars. The line if ((scores[i] > min_conf_threshold) and (scores[i] <= 1.0)): is the culprit—this means scores[i] is an array, not a single value.
What's going wrong:
You're likely grabbing an auxiliary output (like raw_detection_scores, a 3D array) instead of the actual detection_scores (a 2D array of shape [1, max_detections]).
Fix steps:
- First, correctly identify which TFLite output corresponds to
detection_scores(use the comparison method above). - Remove the batch dimension from the scores tensor to get a 1D array:
# After getting outputs from TFLite interpreter scores = outputs[correct_scores_index][0] # Extracts the first (only) batch element, making it 1D - Iterate only over valid detections (up to
num_detections):num_detections = int(outputs[num_detections_index][0][0]) for i in range(num_detections): score = scores[i] if score > min_conf_threshold and score <= 1.0: # Process your detection here (boxes, classes, etc.) pass
By squeezing the batch dimension, scores[i] becomes a scalar value, which fixes the boolean condition ambiguity.
内容的提问来源于stack exchange,提问作者A.Dumas

