TensorFlow定点量化在嵌入式DSP上的底层实现疑问
Hey Anton, great question—let’s break this down step by step since non-symmetric quantization can feel counterintuitive when you’re used to symmetric fixed-point shifts. The key is to reframe the floating-point multiply-accumulate (MAC) operation entirely as integer math (leveraging your DSP’s 9-bit MACs) plus a tiny bit of offline precomputation and final floating-point scaling.
First, Let’s Formalize Non-Symmetric Quantization
For any layer, we map floating-point values (weights w or inputs x) to 8-bit integers using a linear scaling with an offset:
- For weights:
w = s_w * (q_w - zp_w)s_w = (w_max - w_min) / 255(scales the full weight range to 0-255)zp_w = round(-w_min / s_w)(the 8-bit integer that maps to floating-point 0, or shifts the range sow_minmaps toq_w=0andw_maxmaps toq_w=255)
- For inputs:
x = s_x * (q_x - zp_x)(same logic with input-specifics_x,zp_x,x_min,x_max)
Rewrite the MAC Operation for Integer Math
The core inference operation (e.g., convolution/full connect) is y = sum(w_i * x_i) + b. Substitute the quantized forms of w and x:
sum(w_i * x_i) = sum( s_w*(q_wi - zp_w) * s_x*(q_xi - zp_x) ) = s_w*s_x * sum( (q_wi - zp_w)*(q_xi - zp_x) )
Wait, that’s not the whole story—if your non-symmetric range doesn’t include 0, we need to expand the product to capture all terms, but crucially every sum inside uses only integers:
sum(w_i*x_i) = s_w*s_x * sum_offset + s_w*w_min * sum_x_offset + s_x*x_min * sum_w_offset + N * w_min*x_min
Where:
sum_offset = sum( (q_wi - zp_w) * (q_xi - zp_x) )→ pure integer MAC, perfect for your DSPsum_x_offset = sum(q_xi - zp_x) = sum(q_xi) - N*zp_x→ integer sum of input offsets (easy to compute on-chip)sum_w_offset = sum(q_wi - zp_w)→ precomputed offline (weights are fixed, so calculate once and store)N= number of MACs per operation (e.g., convolution kernel size, fixed per layer)s_w*w_min,s_x*x_min,N*w_min*x_min→ precomputed floating-point constants (offline, once per layer)
Adapt This to Your DSP’s 9-Bit MACs
Your DSP takes 9-bit inputs, which is a perfect fit:
(q_wi - zp_w)and(q_xi - zp_x)range from-255to255(sinceq_w/q_xare 0-255,zp_w/zp_xare 0-255 integers). This exactly fits in a signed 9-bit integer (-256 to 255).- Preprocess your quantized weights offline: convert each 8-bit
q_wito its signed 9-bit offsetq_wi - zp_wand store these values. No need to compute this on-chip during inference. - For inputs: quantize to 8-bit
q_xi, then computeq_xi - zp_xon-chip (a simple integer subtraction) to get the 9-bit input for the DSP.
Step-by-Step Inference Workflow
Offline Prep (Per Layer):
- Calculate
w_min,w_max,s_w,zp_wfor weights; same for inputs using calibration data. - Convert weights to 8-bit
q_w, then precomputesum_w_offsetand store the 9-bit offset valuesq_w - zp_w. - Precompute floating-point constants:
s_total = s_w*s_x,term1 = s_w*w_min,term2 = s_x*x_min,const_term = N*w_min*x_min.
- Calculate
On-Chip Inference:
- Quantize input to 8-bit
q_x, compute 9-bit input offsetsq_x - zp_x. - Use your DSP’s 168-cycle MAC to compute
sum_offset(all integer operations). - Compute
sum_x_offset = sum(q_x) - N*zp_x(integer sum of input values minus a precomputed constant). - Calculate the final floating-point result:
result = s_total * sum_offset + term1 * sum_x_offset + term2 * sum_w_offset + const_term + b - (Optional) Apply activation (e.g., ReLU) and quantize the result for the next layer.
- Quantize input to 8-bit
Key Notes
- Overflow Safety: Each individual MAC result ranges from
(-255)*(-255)=65025to255*255=65025. For 168 MACs, the totalsum_offsetranges from ~-10.9M to ~10.9M, which fits easily in a 32-bit integer accumulator (your DSP almost certainly has this). - Accuracy: This approach preserves the same minimal accuracy loss as standard 8-bit quantization, since we’re just rearranging the math rather than truncating or approximating differently.
- DSP Efficiency: 99% of the compute work happens in your DSP’s fast, low-power MACs—only the final scaling step uses floating-point operations (a tiny overhead compared to the MAC workload).
内容的提问来源于stack exchange,提问作者Anton Krug

