You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ML.NET调用HuggingFace RMBG-2.0 ONNX模型结果异常求助

排查ML.NET运行RMBG-2.0 ONNX模型推理结果不准确的问题

我尝试使用ML.NET运行HuggingFace的RMBG-2.0图像分割ONNX模型进行推理,代码已可编译输出结果,但结果与HuggingFace演示页面差异极大。已知代码存在优化空间(如替换GetPixel/SetPixel),且未缩放回原尺寸导致比例错误,但当前优先解决背景去除准确性问题。

我的代码

public static void RemoveGreenBackgroundAI2(string imagePath, string outputfile)
{
    string modelPath = Path.Combine( Application.StartupPath,"ONNX","model.onnx"); 
    MLContext mlContext = new MLContext();

    var imageData = new ImageInputData
    {
        Image = MLImage.CreateFromFile (imagePath)
    };


    var imageDataView = mlContext.Data.LoadFromEnumerable(new[] { imageData });
   
   var pipeline = mlContext.Transforms.ResizeImages(
                        outputColumnName: "input",
                        imageWidth: 1024,
                        imageHeight: 1024,
                        inputColumnName: nameof(ImageInputData.Image))
                  .Append(mlContext.Transforms.ExtractPixels(
                        outputColumnName: "out1",
                        inputColumnName: "input",
                        interleavePixelColors: true,
                        scaleImage: 1f / 255f,
                        offsetImage: 0,
                        outputAsFloatArray: true))
                    .Append(mlContext.Transforms.CustomMapping<CustomMappingInput, CustomMappingOutput>( 
                       mapAction: (input, output) =>
                        {
                            output.pixel_values = new float[input.out1.Length];
                            for (int i = 0; i < input.out1.Length; i += 3)
                            {
                                // R
                                output.pixel_values[i] = (input.out1[i] - 0.485f) / 0.229f;

                                //G
                                output.pixel_values[i + 1] = (input.out1[i + 1] - 0.456f) / 0.224f;

                                //B
                                output.pixel_values[i + 2] = (input.out1[i + 2] - 0.406f) / 0.225f;
                            }
                        }, contractName: null))
                  .Append(mlContext.Transforms.ApplyOnnxModel(
                        modelFile: modelPath,
                        outputColumnNames: new[] { "alphas" },
                        inputColumnNames: new[] { "pixel_values" },
                        shapeDictionary: new Dictionary<string, int[]>
                        {
                            { "pixel_values", new[] { 1, 3, 1024, 1024 } }

                        },
                        fallbackToCpu:true,
                        gpuDeviceId:null
                        ));

    
    var model = pipeline.Fit(imageDataView);
    var predictionEngine = mlContext.Model.CreatePredictionEngine<ImageInputData, ModelOutput>(model);
    var prediction = predictionEngine.Predict(imageData);
    ApplyMaskAndSaveImage(imagePath, prediction, outputfile);

}

public static void ApplyMaskAndSaveImage(string originalImagepath, ModelOutput prediction, string outputPath)
{
    int width = 1024;
    int height = 1024;
    float[] outputData = prediction.Output;

    Bitmap originalImage = (Bitmap)Bitmap.FromFile(originalImagepath);
    int originalWidth = originalImage.Width;
    int originalHeight = originalImage.Height;

    Bitmap resizedImage = new Bitmap(originalImage, new System.Drawing.Size(width, height));
    Bitmap outputImage = new Bitmap(width, height, PixelFormat.Format32bppArgb);

    for (int y = 0; y < height; y++)
    {
        for (int x = 0; x < width; x++)
        {
            float maskValue = outputData[y * width + x];
            float threshold = 0.5f;
            byte alpha = maskValue >= threshold ? (byte)255 : (byte)0;
            Color pixelColor = resizedImage.GetPixel(x, y);
            Color newColor = Color.FromArgb(alpha, pixelColor.R, pixelColor.G, pixelColor.B);
            outputImage.SetPixel(x, y, newColor);
        }
    }      
    outputImage.Save(outputPath, ImageFormat.Png);
}

public class ModelOutput
{
    [ColumnName("alphas")]
    [VectorType(1, 1, 1024, 1024)]
    public float[] Output { get; set; }
}
public class ImageInputData
{
    [ColumnName("Image")]
    [ImageType(1024, 1024)]
    public MLImage Image { get; set; }
}
public class CustomMappingInput
{
    [VectorType(3, 1024, 1024)]
    public float[] out1 { get; set; }
}
public class CustomMappingOutput
{
    [VectorType(3, 1024, 1024)]
    public float[] pixel_values { get; set; } 
}

排查建议

1. 修正图像预处理的通道顺序与数据排布

  • 当前ExtractPixels设置了interleavePixelColors: true,输出的是NHWC格式(每个像素的R、G、B值连续存储),但CustomMappingInput标注的VectorType(3,1024,1024)是NCHW格式的维度,两者不匹配,会导致数据解析错误。
  • 解决方案:将ExtractPixels的interleavePixelColors改为false,此时输出为NCHW格式(所有R像素→所有G像素→所有B像素),与你标注的VectorType一致,后续归一化逻辑可以直接按通道处理。
  • 同时确认官方模型的通道顺序:若模型训练时用的是BGR输入,需在归一化前交换R和B通道的位置。

2. 验证ONNX模型的输入输出维度映射

  • 检查ApplyOnnxModel中输入输出的维度是否与模型要求完全一致:
    • 输入pixel_values的形状[1,3,1024,1024]是否符合模型要求(多数图像分割模型为NCHW);
    • 输出alphas的形状是否为[1,1,1024,1024],若模型输出是NHWC格式,需调整ModelOutput的VectorType为[1,1024,1024,1],并修改掩码索引逻辑。

3. 可视化模型输出的掩码数据

  • 在ApplyMaskAndSaveImage中,先将outputData直接转为灰度图保存(将float值映射到0-255的字节范围),查看掩码是否符合预期。如果灰度图完全混乱,说明输出维度解析错误,需调整ModelOutput的VectorType或索引逻辑。

4. 对比官方预处理后的输入数据

  • 将ML.NET预处理后的pixel_values导出为数组,与官方Python代码预处理后的输入数组对比数值。若差异较大,说明预处理步骤(缩放、归一化、通道顺序)存在错误。

5. 验证ONNX模型的兼容性

  • 用ONNX Runtime直接加载模型进行推理(脱离ML.NET),传入官方预处理后的输入数据,看是否能得到正确掩码。如果结果正确,说明问题出在ML.NET的pipeline配置上;如果结果仍错误,需确认ONNX模型的版本与导出方式是否正确。

内容的提问来源于stack exchange,提问作者alepee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 15:05:58