ONNX模型Opset 9升级至11后量化失败的替代方案咨询
Got it, let's work through this quantization problem you're facing with your Pictor-PPE model. The root issue here is that the automatic opset upgrade to 11 broke the deprecated Upsample operator, which is why your quantized model won't load. Here are four solid alternative approaches to get a smaller, working quantized model:
1. Manually Fix the Upsample Operator Before Quantization
The automatic opset upgrade didn't properly replace the deprecated Upsample with its recommended replacement (Resize) for opset 11+. Let's fix this manually:
import onnx from onnx.version_converter import convert_version # Load and upgrade the original model to opset 11 model = onnx.load("path/to/model.onnx") model = convert_version(model, 11) graph = model.graph nodes_to_remove = [] new_nodes = [] for node in graph.node: if node.op_type == "Upsample": # Create a Resize node to replace Upsample resize_node = onnx.helper.make_node( "Resize", inputs=[node.input[0], "", ""], outputs=node.output, name=node.name.replace("Upsample", "Resize") ) # Convert Upsample's scale attribute to a Resize-compatible scales tensor scale_attr = node.attribute[0] scales = onnx.helper.make_tensor( name=f"{node.name}_scales", data_type=onnx.TensorProto.FLOAT, dims=[4], # Match NHWC input shape: [1, scale_h, scale_w, 1] vals=[1.0, scale_attr.floats[0], scale_attr.floats[1], 1.0] ) graph.initializer.append(scales) resize_node.input[1] = scales.name new_nodes.append(resize_node) nodes_to_remove.append(node) else: new_nodes.append(node) # Update the graph with new nodes and validate graph.node[:] = new_nodes for node in nodes_to_remove: graph.node.remove(node) onnx.checker.check_model(model) onnx.save(model, "path/to/model_opset11_fixed.onnx") # Now run dynamic quantization on the fixed model from onnxruntime.quantization import quantize_dynamic, QuantType quantize_dynamic( "path/to/model_opset11_fixed.onnx", "path/to/model_quant_fixed.onnx", weight_type=QuantType.QUInt8 )
2. Quantize in TensorFlow First, Then Convert to ONNX
Since your original model is in TensorFlow, doing quantization directly in TF avoids ONNX operator compatibility headaches entirely. Try these options:
Dynamic Quantization (Quick, CPU-friendly)
import tensorflow as tf from tensorflow.keras.layers import Input # Reload your original TF model input_tensor = Input(shape=(input_shape[0], input_shape[1], 3)) num_out_filters = (num_anchors//3) * (5 + num_classes) model = yolo_body(input_tensor, num_out_filters) weight_path = 'ONNX_demo/models/pictor-ppe-v302-a1-yolo-v3-weights.h5' model.load_weights(weight_path) # Apply dynamic quantization quantized_model = tf.keras.models.quantize_model(model) quantized_model.compile(optimizer='adam', loss='categorical_crossentropy') # Save and convert to ONNX tf.saved_model.save(quantized_model, "ONNX_demo/models/quantized_save_model") !python3 -m tf2onnx.convert --saved-model "ONNX_demo/models/quantized_save_model" --output "ONNX_demo/models/model_quant_tf.onnx"
Static Quantization (Better Accuracy, Needs Calibration Data)
For higher precision, use static quantization with a calibration dataset:
# Reload model (same as above) # Define calibration data generator def calibration_dataset(): # Replace with real input data matching your model's shape for _ in range(100): yield [tf.random.normal([1, input_shape[0], input_shape[1], 3])] # Convert to quantized TFLite model converter = tf.lite.TFLiteConverter.from_keras_model(model) converter.optimizations = [tf.lite.Optimize.DEFAULT] converter.representative_dataset = calibration_dataset converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8] converter.inference_input_type = tf.uint8 converter.inference_output_type = tf.uint8 tflite_quant_model = converter.convert() with open("ONNX_demo/models/model_quant.tflite", "wb") as f: f.write(tflite_quant_model) # Convert TFLite to ONNX !python3 -m tf2onnx.convert --tflite "ONNX_demo/models/model_quant.tflite" --output "ONNX_demo/models/model_quant_tflite.onnx"
3. Simplify the ONNX Model Before Quantization
ONNX Simplifier automatically cleans up redundant operators and replaces deprecated ones (like Upsample with Resize) which can fix your quantization error:
- Install the tool:
pip install onnx-simplifier
- Simplify your model:
python -m onnxsim path/to/model.onnx path/to/model_simplified.onnx
- Quantize the simplified model:
from onnxruntime.quantization import quantize_dynamic, QuantType quantize_dynamic( "path/to/model_simplified.onnx", "path/to/model_quant_simplified.onnx", weight_type=QuantType.QUInt8 )
4. Use TensorRT for Quantization (GPU Deployment)
If you're deploying on an NVIDIA GPU, TensorRT's INT8 quantization reduces model size and boosts inference speed:
import tensorrt as trt TRT_LOGGER = trt.Logger(trt.Logger.WARNING) builder = trt.Builder(TRT_LOGGER) network = builder.create_network(1 << int(trt.NetworkDefinitionCreationFlag.EXPLICIT_BATCH)) parser = trt.OnnxParser(network, TRT_LOGGER) # Load your ONNX model with open("path/to/model.onnx", "rb") as f: parser.parse(f.read()) # Configure INT8 quantization config = builder.create_builder_config() config.set_memory_pool_limit(trt.MemoryPoolType.WORKSPACE, 1 << 30) # 1GB workspace config.set_flag(trt.BuilderFlag.INT8) # Build and save the quantized TensorRT engine engine = builder.build_engine(network, config) with open("path/to/model.trt", "wb") as f: f.write(engine.serialize())
内容的提问来源于stack exchange,提问作者Luis Ramon Ramirez Rodriguez

