使用C# ML.NET调用DepthAnything V2模型生成深度图失败求助
DepthAnything V2 C#代码生成异常深度图的问题排查与修复
问题描述
尝试使用DepthAnything V2模型生成深度图,编写了C#代码后生成异常图像,直接用[x,y]索引输出张量时触发越界错误,输出张量大小为268324(对应518×518)。代码如下:
using System.Drawing; using System.Text.Json; using Microsoft.ML; using Microsoft.ML.OnnxRuntime; using Microsoft.ML.OnnxRuntime.Tensors; public static class Program { static Tensor<float> LoadImage(string imagePath) { var image = new Bitmap(imagePath); var inputTensor = new DenseTensor<float>(new[] { 1, 3, image.Height, image.Width }); for (int y = 0; y < image.Height; y++) { for (int x = 0; x < image.Width; x++) { Color pixel = image.GetPixel(x, y); inputTensor[0, 0, y, x] = pixel.R / 255.0f; inputTensor[0, 1, y, x] = pixel.G / 255.0f; inputTensor[0, 2, y, x] = pixel.B / 255.0f; } } return inputTensor; } public static void Main() { Tensor<float> image = LoadImage("demo01.jpg"); MLContext mlContext = new MLContext(); InferenceSession inferenceSession = new InferenceSession("depth_anything_v2_vits.onnx"); Console.WriteLine(JsonSerializer.Serialize(inferenceSession.InputNames)); var inputs = new NamedOnnxValue[] { NamedOnnxValue.CreateFromTensor("l_x_", image) }; Tensor<float> output = inferenceSession.Run(inputs).First().AsTensor<float>(); Bitmap outImage = new Bitmap(518, 518); Console.WriteLine(output.Length); for(int y = 0; y < 518; y++) for(int x = 0; x < 518; x++) { int r = (int)Math.Floor(output.ElementAt(x * y) * 50); outImage.SetPixel(x, y, Color.FromArgb(255, r, 0, 0)); } outImage.Save("out.png"); } }
问题排查与修复
1. 输入图像预处理不符合模型要求
DepthAnything V2要求输入图像必须是固定尺寸(如518×518),且需遵循特定归一化规则:
- 代码直接读取原图尺寸输入,若原图非518×518,会导致模型输出异常;
- 仅将像素值除以255不符合模型规范,需用均值
[0.485, 0.456, 0.406]和标准差[0.229, 0.224, 0.225]做归一化。
修复后的预处理代码:
static Tensor<float> LoadImage(string imagePath) { // 模型要求的固定输入尺寸 int targetWidth = 518; int targetHeight = 518; using var image = new Bitmap(imagePath); // 先将图像缩放到目标尺寸 using var resizedImage = new Bitmap(image, targetWidth, targetHeight); var inputTensor = new DenseTensor<float>(new[] { 1, 3, targetHeight, targetWidth }); // 模型标准归一化参数 float[] mean = new float[] { 0.485f, 0.456f, 0.406f }; float[] std = new float[] { 0.229f, 0.224f, 0.225f }; for (int y = 0; y < targetHeight; y++) { for (int x = 0; x < targetWidth; x++) { Color pixel = resizedImage.GetPixel(x, y); // 执行归一化计算 inputTensor[0, 0, y, x] = (pixel.R / 255.0f - mean[0]) / std[0]; inputTensor[0, 1, y, x] = (pixel.G / 255.0f - mean[1]) / std[1]; inputTensor[0, 2, y, x] = (pixel.B / 255.0f - mean[2]) / std[2]; } } return inputTensor; }
2. 输出张量索引逻辑完全错误
输出张量是形状为[1, 1, 518, 518]的4维张量,原代码用x * y计算索引的逻辑存在严重问题:
- 当x或y为0时,索引始终为0,会重复读取同一个元素;
- 索引计算不符合张量的行优先存储规则,导致图像像素映射完全混乱。
正确的索引方式有两种:
- 直接使用4维索引:
output[0, 0, y, x] - 计算线性索引:
y * 518 + x(行优先存储,先遍历x再遍历y)
修复后的图像生成代码:
Bitmap outImage = new Bitmap(518, 518); // 获取深度值的范围,用于归一化到0-255区间 float minDepth = output.Min(); float maxDepth = output.Max(); for(int y = 0; y < 518; y++) { for(int x = 0; x < 518; x++) { // 正确读取深度值 float depthValue = output[0, 0, y, x]; // 归一化到0-255,反转后近景(深度值大)更亮 int grayValue = (int)Math.Round(255 * (maxDepth - depthValue) / (maxDepth - minDepth)); grayValue = Math.Clamp(grayValue, 0, 255); // 防止数值溢出 outImage.SetPixel(x, y, Color.FromArgb(255, grayValue, grayValue, grayValue)); } } outImage.Save("out.png");
3. 输入张量名称需确认匹配
原代码使用"l_x_"作为输入名称,建议通过以下代码确认模型实际输入名称和形状:
foreach (var inputMeta in inferenceSession.InputMetadata) { Console.WriteLine($"输入名称: {inputMeta.Key}, 形状: {string.Join(",", inputMeta.Value.Dimensions)}"); }
确保输入张量的名称和形状与模型要求完全一致。
内容的提问来源于stack exchange,提问作者user21053170
相关产品推荐
相关产品推荐

