You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ONNX模型Opset 9升级至11后量化失败的替代方案咨询

Got it, let's work through this quantization problem you're facing with your Pictor-PPE model. The root issue here is that the automatic opset upgrade to 11 broke the deprecated Upsample operator, which is why your quantized model won't load. Here are four solid alternative approaches to get a smaller, working quantized model:

1. Manually Fix the Upsample Operator Before Quantization

The automatic opset upgrade didn't properly replace the deprecated Upsample with its recommended replacement (Resize) for opset 11+. Let's fix this manually:

import onnx
from onnx.version_converter import convert_version

# Load and upgrade the original model to opset 11
model = onnx.load("path/to/model.onnx")
model = convert_version(model, 11)

graph = model.graph
nodes_to_remove = []
new_nodes = []

for node in graph.node:
    if node.op_type == "Upsample":
        # Create a Resize node to replace Upsample
        resize_node = onnx.helper.make_node(
            "Resize",
            inputs=[node.input[0], "", ""],
            outputs=node.output,
            name=node.name.replace("Upsample", "Resize")
        )
        
        # Convert Upsample's scale attribute to a Resize-compatible scales tensor
        scale_attr = node.attribute[0]
        scales = onnx.helper.make_tensor(
            name=f"{node.name}_scales",
            data_type=onnx.TensorProto.FLOAT,
            dims=[4],  # Match NHWC input shape: [1, scale_h, scale_w, 1]
            vals=[1.0, scale_attr.floats[0], scale_attr.floats[1], 1.0]
        )
        graph.initializer.append(scales)
        resize_node.input[1] = scales.name
        
        new_nodes.append(resize_node)
        nodes_to_remove.append(node)
    else:
        new_nodes.append(node)

# Update the graph with new nodes and validate
graph.node[:] = new_nodes
for node in nodes_to_remove:
    graph.node.remove(node)

onnx.checker.check_model(model)
onnx.save(model, "path/to/model_opset11_fixed.onnx")

# Now run dynamic quantization on the fixed model
from onnxruntime.quantization import quantize_dynamic, QuantType
quantize_dynamic(
    "path/to/model_opset11_fixed.onnx",
    "path/to/model_quant_fixed.onnx",
    weight_type=QuantType.QUInt8
)

2. Quantize in TensorFlow First, Then Convert to ONNX

Since your original model is in TensorFlow, doing quantization directly in TF avoids ONNX operator compatibility headaches entirely. Try these options:

Dynamic Quantization (Quick, CPU-friendly)

import tensorflow as tf
from tensorflow.keras.layers import Input

# Reload your original TF model
input_tensor = Input(shape=(input_shape[0], input_shape[1], 3))
num_out_filters = (num_anchors//3) * (5 + num_classes)
model = yolo_body(input_tensor, num_out_filters)
weight_path = 'ONNX_demo/models/pictor-ppe-v302-a1-yolo-v3-weights.h5'
model.load_weights(weight_path)

# Apply dynamic quantization
quantized_model = tf.keras.models.quantize_model(model)
quantized_model.compile(optimizer='adam', loss='categorical_crossentropy')

# Save and convert to ONNX
tf.saved_model.save(quantized_model, "ONNX_demo/models/quantized_save_model")
!python3 -m tf2onnx.convert --saved-model "ONNX_demo/models/quantized_save_model" --output "ONNX_demo/models/model_quant_tf.onnx"

Static Quantization (Better Accuracy, Needs Calibration Data)

For higher precision, use static quantization with a calibration dataset:

# Reload model (same as above)

# Define calibration data generator
def calibration_dataset():
    # Replace with real input data matching your model's shape
    for _ in range(100):
        yield [tf.random.normal([1, input_shape[0], input_shape[1], 3])]

# Convert to quantized TFLite model
converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = calibration_dataset
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.uint8
converter.inference_output_type = tf.uint8

tflite_quant_model = converter.convert()
with open("ONNX_demo/models/model_quant.tflite", "wb") as f:
    f.write(tflite_quant_model)

# Convert TFLite to ONNX
!python3 -m tf2onnx.convert --tflite "ONNX_demo/models/model_quant.tflite" --output "ONNX_demo/models/model_quant_tflite.onnx"

3. Simplify the ONNX Model Before Quantization

ONNX Simplifier automatically cleans up redundant operators and replaces deprecated ones (like Upsample with Resize) which can fix your quantization error:

  1. Install the tool:
pip install onnx-simplifier
  1. Simplify your model:
python -m onnxsim path/to/model.onnx path/to/model_simplified.onnx
  1. Quantize the simplified model:
from onnxruntime.quantization import quantize_dynamic, QuantType
quantize_dynamic(
    "path/to/model_simplified.onnx",
    "path/to/model_quant_simplified.onnx",
    weight_type=QuantType.QUInt8
)

4. Use TensorRT for Quantization (GPU Deployment)

If you're deploying on an NVIDIA GPU, TensorRT's INT8 quantization reduces model size and boosts inference speed:

import tensorrt as trt

TRT_LOGGER = trt.Logger(trt.Logger.WARNING)
builder = trt.Builder(TRT_LOGGER)
network = builder.create_network(1 << int(trt.NetworkDefinitionCreationFlag.EXPLICIT_BATCH))
parser = trt.OnnxParser(network, TRT_LOGGER)

# Load your ONNX model
with open("path/to/model.onnx", "rb") as f:
    parser.parse(f.read())

# Configure INT8 quantization
config = builder.create_builder_config()
config.set_memory_pool_limit(trt.MemoryPoolType.WORKSPACE, 1 << 30)  # 1GB workspace
config.set_flag(trt.BuilderFlag.INT8)

# Build and save the quantized TensorRT engine
engine = builder.build_engine(network, config)
with open("path/to/model.trt", "wb") as f:
    f.write(engine.serialize())

内容的提问来源于stack exchange,提问作者Luis Ramon Ramirez Rodriguez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 00:32:27