You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

iOS平台TensorFlowLite版YOLOv4输出解析问题咨询

Hey there! Let's tackle your two questions step by step— I’ve messed around with YOLOv4 + TensorFlow Lite on iOS before, so this hits close to home.

1. What's the purpose of the 26x26 output tensor?

Absolutely, that second output is for finer-grained object detection. YOLOv4 uses a multi-scale detection pipeline (part of its "feature pyramid" design) where different grid sizes target different object sizes:

  • The 13x13 grid is coarser, optimized for detecting large objects (since each cell covers more pixels in the original image).
  • The 26x26 grid is more dense, so it can pick up smaller or medium-sized objects that the 13x13 grid might miss (each cell covers a smaller area, allowing it to capture finer details).

Some YOLOv4 variants even include a 52x52 grid for tiny objects, but your model stops at two— that’s totally normal for many deployment setups.

2. Why aren't the anchor values in the 0-1 range?

Your suspicion about preprocessing is valid, but first there’s a bigger piece you’re missing: YOLOv4’s output are raw logits, not final probabilities/coordinates. They need post-processing to get into the 0-1 range or meaningful bounding box values. Let’s break this down:

First, fix the order of your output values

You mentioned you expect the first 5 values to be "confidence, BBoxX, BBoxY, BBoxWidth, BBoxHeight"— that’s reversed! The standard YOLOv4 anchor output order is:
[tx, ty, tw, th, confidence_logit, class_1_logit, ..., class_80_logit]

  • tx, ty: Offset from the top-left corner of the current grid cell (raw logits)
  • tw, th: Scaling factor for the pre-defined anchor box size (raw logits)
  • confidence_logit: How confident the model is that this anchor contains an object (raw logit)
  • class_N_logit: Probability logits for each of the 80 classes

Post-processing steps to get valid values

You need to apply activation functions and coordinate transformations to these logits:

  1. Sigmoid for offsets, confidence, and class probabilities
    The sigmoid function converts any real number to a value between 0 and 1. You’ll need this for:

    • tx → bx: Relative offset within the grid cell (add the cell’s x index to get absolute grid position)
    • ty → by: Same for y position
    • confidence_logit → final confidence score
    • Each class_N_logit → class probability

    Add this helper function to your code:

    func sigmoid(_ x: Float) -> Float {
        return 1.0 / (1.0 + exp(-x))
    }
    
  2. Exponential function for bounding box width/height
    tw and th are scaling factors for your pre-defined anchor sizes. You’ll multiply the anchor’s width/height by exp(tw) and exp(th) to get the actual bounding box dimensions.

  3. Convert to image-normalized values
    Once you have the grid-based coordinates, divide by the grid size (13 or 26) to get values between 0 and 1 relative to the image. For width/height, divide by your input image size (usually 416x416 for YOLOv4).

Example post-processing for your anchor data

Here’s how to modify your loop to process each anchor correctly:

// First, define your pre-trained anchor sizes (must match what was used to train the model!)
// For YOLOv4, 13x13 anchors are typically [12,16], [19,36], [40,28]
let anchors13: [[Float]] = [[12.0, 16.0], [19.0, 36.0], [40.0, 28.0]]
// 26x26 anchors are usually [36,75], [76,55], [72,146]
let anchors26: [[Float]] = [[36.0, 75.0], [76.0, 55.0], [72.0, 146.0]]

for y in 0..<numberOfCells {
    for x in 0..<numberOfCells {
        let cell = cellAt(x: x, y: y)
        let anchors = anchors(in: cell)
        
        for (anchorIndex, anchor) in anchors.enumerated() {
            // Extract raw logits
            let tx = anchor[0]
            let ty = anchor[1]
            let tw = anchor[2]
            let th = anchor[3]
            let confidenceLogit = anchor[4]
            let classLogits = Array(anchor[5...])
            
            // Process coordinates
            let bx = sigmoid(tx) + Float(x)
            let by = sigmoid(ty) + Float(y)
            // Normalize to 0-1 relative to image
            let bxNorm = bx / Float(numberOfCells)
            let byNorm = by / Float(numberOfCells)
            
            // Process width/height
            let anchorW = anchors13[anchorIndex][0]
            let anchorH = anchors13[anchorIndex][1]
            let bw = anchorW * exp(tw)
            let bh = anchorH * exp(th)
            // Normalize to 0-1 relative to 416x416 input
            let bwNorm = bw / 416.0
            let bhNorm = bh / 416.0
            
            // Process confidence and class probabilities
            let confidence = sigmoid(confidenceLogit)
            let classProbs = classLogits.map { sigmoid($0) }
            
            // Now you have valid 0-1 values!
            print("Cell (\(x),\(y)) Anchor \(anchorIndex):")
            print("Confidence: \(confidence), BBox: (\(bxNorm), \(byNorm), \(bwNorm), \(bhNorm))")
            print("Top class prob: \(classProbs.max() ?? 0.0)")
        }
    }
}

Preprocessing mismatch check

Your hunch about imageMean and imageStd is still worth verifying. YOLOv4 training usually uses one of two preprocessing pipelines:

  1. Normalize pixel values to 0-1 by dividing by 255.0 (no mean/std subtraction/scaling)
  2. Scale pixels to -1 to 1 by dividing by 127.5 and subtracting 1.0

If your training pipeline used the first method but you’re using imageMean=127.5 in your iOS code, the input distribution will be wrong, leading to weird output logits. Double-check your training config to match the preprocessing exactly.


内容的提问来源于stack exchange,提问作者Damian Dudycz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:12:53