You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python与C++用OpenCV归一化图像后数据不一致原因及对齐方法

Python与C++ OpenCV图像归一化结果不一致问题排查与解决

我在使用Python和C++的OpenCV将图像归一化到0-1区间后,发现处理后的图像数据完全不同,目标是让两者预处理结果完全一致。原始图像如下:

原始图像

我的C++代码

void writeImageDataToFile(const cv::Mat& image, const std::string& filePath) {
    std::ofstream outputFile(filePath);
    if (!outputFile.is_open()) {
        std::cout << "Failed to open file for writing: " << filePath << std::endl;
        return;
    }

    for (int channel = 0; channel < image.channels(); ++channel) {
        outputFile << "Channel " << channel << ":" << std::endl;
        for (int row = 0; row < image.rows; ++row) {
            for (int col = 0; col < image.cols; ++col) {
                outputFile << std::fixed << std::setprecision(6) << image.at<cv::Vec3f>(row, col)[channel] << " ";
            }
            outputFile << std::endl;
        }
        outputFile << std::endl;
    }

    outputFile.close();
}

int cntt = -1;
for (const auto& directory : directories) {
    std::vector<std::string> imagePaths = GetImagePaths(directory);

    for (const auto& imagePath : imagePaths) {
        cv::Mat image = cv::imread(imagePath);
        if (image.empty()) {
            std::cout << "Failed to read image: " << imagePath << std::endl;
            continue;
        }

        cntt++;
        cv::cvtColor(image, image, cv::COLOR_BGR2RGB); 

        cv::Mat resizedImage;
        cv::resize(image, resizedImage, cv::Size(32, 32));
        
        cv::Mat normalizedImage;
        resizedImage.convertTo(normalizedImage, CV_32F);

        cv::Scalar new_mean(0.9966, 0.9973, 0.9974);
        cv::Scalar new_std(0.0102, 0.0078, 0.0074);

        normalizedImage /= 255.0;
        for (int i = 0; i < 3; i++) {
            normalizedImage.row(i) = (normalizedImage.row(i) - new_mean[i]) / new_std[i];
        }
        
        std::string outputFilePath = "normalized_image_data" + std::to_string(cntt) + ".txt";
        writeImageDataToFile(normalizedImage, outputFilePath);
    }
}

我的Python代码

import os
import cv2
import numpy as np

def write_image_data_to_file(image, file_path):
    with open(file_path, 'w') as output_file:
        for channel in range(image.shape[0]):
            output_file.write(f"Channel {channel}:\n")
            for row in range(image.shape[1]):
                for col in range(image.shape[2]):
                    output_file.write(f"{image[channel, row, col]:.6f} ")
                output_file.write("\n")
            output_file.write("\n")


input_folder = 'input0'  # Change to your input folder name
output_folder = 'output_folder3'  # Change to your desired output folder name
new_mean = [0.9966, 0.9973, 0.9974]
new_std = [0.0102, 0.0078, 0.0074]


if not os.path.exists(output_folder):
    os.makedirs(output_folder)


for filename in os.listdir(input_folder):
    if filename.endswith(".jpg"):
        image_path = os.path.join(input_folder, filename)

        image = cv2.imread(image_path)

        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)
        image = cv2.resize(image, (32, 32))
        image = np.transpose(image, (2, 0, 1))

        image = (image / 255.0).astype(np.float32)
        for i in range(3):
            image[i] = (image[i] - new_mean[i]) / new_std[i]

        output_file_path = os.path.join(output_folder, filename.replace(".jpg", ".txt"))

        write_image_data_to_file(image, output_file_path)

输出差异

C++代码输出示例(normalized_image_data0.txt)

Channel 0:
0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 ...(32 count)
0.346153 0.346153  ...
0.351349 0.351349 ...
1.000000 1.000000 ...
1.000000 1.000000 ...
...

Channel 1:
0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 0.333336 ...
0.346153 0.346153 ...
...
1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 0.980392 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 
...

Channel 2:
0.333336 0.333336 ...
0.346153 0.346153 ...
...
1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 0.992157 0.988235 0.980392 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 1.000000 
...

Python代码输出示例

Channel 0:
0.333336 0.333336 ...
0.333336 0.333336 ...
...

Channel 1:
0.346150 0.346150 ...
0.346150 0.346150 ...
...
Channel 2:
0.351353 0.351353 ...
0.351353 0.351353 ...
...

问题原因

核心错误是C++代码中混淆了行和通道的操作:

  • Python代码通过np.transpose(image, (2, 0, 1))将图像转为(通道数, 高度, 宽度)的维度顺序,后续循环image[i]直接操作第i个通道的所有像素。
  • 而C++的cv::Mat默认存储顺序是(高度, 宽度, 通道数),每个像素的三个通道连续存储。但你在标准化步骤中错误使用normalizedImage.row(i)来获取第i个通道——row(i)实际是取图像的第i行像素,而非第i个通道,导致你把行数据当成通道数据做了标准化,完全打乱了数据维度,这是结果完全不同的根本原因。

修正后的C++代码

将标准化部分的代码改为按通道分割处理,修正维度操作:

// ... 前面的读取、转色、resize、convertTo、除以255的代码不变 ...

cv::Scalar new_mean(0.9966, 0.9973, 0.9974);
cv::Scalar new_std(0.0102, 0.0078, 0.0074);

normalizedImage /= 255.0;

// 分割为单个通道
std::vector<cv::Mat> channels(3);
cv::split(normalizedImage, channels);

// 对每个通道执行标准化
for (int i = 0; i < 3; ++i) {
    channels[i] = (channels[i] - new_mean[i]) / new_std[i];
}

// 合并通道回原格式
cv::merge(channels, normalizedImage);

// ... 后续写入文件的代码不变 ...

额外注意事项

  1. 确保cv::resize的插值方式与Python一致:两者默认都是INTER_LINEAR,如果需要严格一致可以显式指定cv::resize(image, resizedImage, cv::Size(32, 32), 0, 0, cv::INTER_LINEAR)。
  2. 浮点数精度差异:C++的float和Python的np.float32精度一致,写入时的std::setprecision(6)与Python的.6f格式匹配,不会导致明显差异。

修正后,C++与Python的预处理结果将完全一致。

内容的提问来源于stack exchange,提问作者kireirain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 15:18:11