iOS平台TensorFlowLite版YOLOv4输出解析问题咨询
Hey there! Let's tackle your two questions step by step— I’ve messed around with YOLOv4 + TensorFlow Lite on iOS before, so this hits close to home.
Absolutely, that second output is for finer-grained object detection. YOLOv4 uses a multi-scale detection pipeline (part of its "feature pyramid" design) where different grid sizes target different object sizes:
- The 13x13 grid is coarser, optimized for detecting large objects (since each cell covers more pixels in the original image).
- The 26x26 grid is more dense, so it can pick up smaller or medium-sized objects that the 13x13 grid might miss (each cell covers a smaller area, allowing it to capture finer details).
Some YOLOv4 variants even include a 52x52 grid for tiny objects, but your model stops at two— that’s totally normal for many deployment setups.
Your suspicion about preprocessing is valid, but first there’s a bigger piece you’re missing: YOLOv4’s output are raw logits, not final probabilities/coordinates. They need post-processing to get into the 0-1 range or meaningful bounding box values. Let’s break this down:
First, fix the order of your output values
You mentioned you expect the first 5 values to be "confidence, BBoxX, BBoxY, BBoxWidth, BBoxHeight"— that’s reversed! The standard YOLOv4 anchor output order is:[tx, ty, tw, th, confidence_logit, class_1_logit, ..., class_80_logit]
tx,ty: Offset from the top-left corner of the current grid cell (raw logits)tw,th: Scaling factor for the pre-defined anchor box size (raw logits)confidence_logit: How confident the model is that this anchor contains an object (raw logit)class_N_logit: Probability logits for each of the 80 classes
Post-processing steps to get valid values
You need to apply activation functions and coordinate transformations to these logits:
Sigmoid for offsets, confidence, and class probabilities
The sigmoid function converts any real number to a value between 0 and 1. You’ll need this for:tx→bx: Relative offset within the grid cell (add the cell’s x index to get absolute grid position)ty→by: Same for y positionconfidence_logit→ final confidence score- Each
class_N_logit→ class probability
Add this helper function to your code:
func sigmoid(_ x: Float) -> Float { return 1.0 / (1.0 + exp(-x)) }Exponential function for bounding box width/height
twandthare scaling factors for your pre-defined anchor sizes. You’ll multiply the anchor’s width/height byexp(tw)andexp(th)to get the actual bounding box dimensions.Convert to image-normalized values
Once you have the grid-based coordinates, divide by the grid size (13 or 26) to get values between 0 and 1 relative to the image. For width/height, divide by your input image size (usually 416x416 for YOLOv4).
Example post-processing for your anchor data
Here’s how to modify your loop to process each anchor correctly:
// First, define your pre-trained anchor sizes (must match what was used to train the model!) // For YOLOv4, 13x13 anchors are typically [12,16], [19,36], [40,28] let anchors13: [[Float]] = [[12.0, 16.0], [19.0, 36.0], [40.0, 28.0]] // 26x26 anchors are usually [36,75], [76,55], [72,146] let anchors26: [[Float]] = [[36.0, 75.0], [76.0, 55.0], [72.0, 146.0]] for y in 0..<numberOfCells { for x in 0..<numberOfCells { let cell = cellAt(x: x, y: y) let anchors = anchors(in: cell) for (anchorIndex, anchor) in anchors.enumerated() { // Extract raw logits let tx = anchor[0] let ty = anchor[1] let tw = anchor[2] let th = anchor[3] let confidenceLogit = anchor[4] let classLogits = Array(anchor[5...]) // Process coordinates let bx = sigmoid(tx) + Float(x) let by = sigmoid(ty) + Float(y) // Normalize to 0-1 relative to image let bxNorm = bx / Float(numberOfCells) let byNorm = by / Float(numberOfCells) // Process width/height let anchorW = anchors13[anchorIndex][0] let anchorH = anchors13[anchorIndex][1] let bw = anchorW * exp(tw) let bh = anchorH * exp(th) // Normalize to 0-1 relative to 416x416 input let bwNorm = bw / 416.0 let bhNorm = bh / 416.0 // Process confidence and class probabilities let confidence = sigmoid(confidenceLogit) let classProbs = classLogits.map { sigmoid($0) } // Now you have valid 0-1 values! print("Cell (\(x),\(y)) Anchor \(anchorIndex):") print("Confidence: \(confidence), BBox: (\(bxNorm), \(byNorm), \(bwNorm), \(bhNorm))") print("Top class prob: \(classProbs.max() ?? 0.0)") } } }
Preprocessing mismatch check
Your hunch about imageMean and imageStd is still worth verifying. YOLOv4 training usually uses one of two preprocessing pipelines:
- Normalize pixel values to 0-1 by dividing by 255.0 (no mean/std subtraction/scaling)
- Scale pixels to -1 to 1 by dividing by 127.5 and subtracting 1.0
If your training pipeline used the first method but you’re using imageMean=127.5 in your iOS code, the input distribution will be wrong, leading to weird output logits. Double-check your training config to match the preprocessing exactly.
内容的提问来源于stack exchange,提问作者Damian Dudycz

