You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++中如何高效写入(9000,9000,4)的3D float向量至文件适配Numpy?

大尺寸图像矩阵的高效读写优化

我的C++程序会生成一个9000×9000的图像矩阵,每个像素包含R、G、B、A四个float类型的颜色值(数值可大于1.0,后续在Python代码中做归一化)。需要将数据保存为文件,方便用Python的numpy.array()或类似工具读取。

目前用CSV格式存储,文件包含8100万行、4列,读写速度慢,文件体积约650MB。

注:每次试验要运行程序多达20次,读写时间和文件体积会累积。

当前C++代码片段

以下是初始化并写入3D向量的代码:

// initializes the vector with data from 'makematrix' class instance
vector<vector<vector<float>>> colorMat = makematrix->getMatrix();

outfile.open("../output/11_14MidRed9k8.csv",std::ios::out);

if (outfile.is_open()) {
    outfile << "r,g,b,a\n"; // writes column labels

    for (unsigned int l=0; l<colorMat.size(); l++) { // 0 to 8999
        for (unsigned int m=0; m<colorMat[0].size(); m++) { // 0 to 8999
            outfile << colorMat[l][m][0] << ',' << colorMat[l][m][1] << ','
                << colorMat[l][m][2] << ',' << colorMat[l][m][3] << '\n';
        }
    }
}

outfile.close();

优化方案建议

1. 改用连续内存+二进制文件存储

当前嵌套vector的内存不连续,遍历和写入效率低。先将数据转为连续的一维数组,再用二进制方式写入,大幅提升读写速度:

C++写入代码:

// 将三维向量转为连续一维数组
vector<float> flatData;
flatData.reserve(colorMat.size() * colorMat[0].size() * 4);
for (const auto& row : colorMat) {
    for (const auto& pixel : row) {
        flatData.insert(flatData.end(), pixel.begin(), pixel.end());
    }
}

// 二进制写入文件,先写入尺寸信息方便Python还原形状
std::ofstream outfile("../output/11_14MidRed9k8.bin", std::ios::binary);
if (outfile.is_open()) {
    int width = static_cast<int>(colorMat[0].size());
    int height = static_cast<int>(colorMat.size());
    int channels = 4;
    // 写入尺寸参数
    outfile.write(reinterpret_cast<const char*>(&width), sizeof(width));
    outfile.write(reinterpret_cast<const char*>(&height), sizeof(height));
    outfile.write(reinterpret_cast<const char*>(&channels), sizeof(channels));
    // 写入原始float数据
    outfile.write(reinterpret_cast<const char*>(flatData.data()), flatData.size() * sizeof(float));
}
outfile.close();

Python读取代码:

import numpy as np

with open("../output/11_14MidRed9k8.bin", "rb") as f:
    # 读取尺寸信息
    width = np.fromfile(f, dtype=np.int32, count=1)[0]
    height = np.fromfile(f, dtype=np.int32, count=1)[0]
    channels = np.fromfile(f, dtype=np.int32, count=1)[0]
    # 读取数据并还原为(9000,9000,4)的矩阵
    data = np.fromfile(f, dtype=np.float32)
    colorMat = data.reshape((height, width, channels))

2. 直接生成numpy原生.npy格式

.npy是numpy的专属二进制格式,自带形状和数据类型信息,Python读取无需额外处理,易用性拉满。可以用轻量级头文件npy.hpp简化C++端的写入逻辑:

C++写入代码(需引入npy.hpp):

#include "npy.hpp"

// 转为连续一维数组
vector<float> flatData;
flatData.reserve(9000 * 9000 * 4);
for (const auto& row : colorMat) {
    for (const auto& pixel : row) {
        flatData.push_back(pixel[0]);
        flatData.push_back(pixel[1]);
        flatData.push_back(pixel[2]);
        flatData.push_back(pixel[3]);
    }
}

// 写入.npy文件,指定形状为(9000,9000,4),数据类型为float32
vector<unsigned long> shape = {9000, 9000, 4};
npy::SaveArrayAsNumpy("../output/11_14MidRed9k8.npy", false, shape.size(), shape.data(), flatData.data());

Python读取代码:

import numpy as np
# 直接读取即可自动还原矩阵形状和数据类型
colorMat = np.load("../output/11_14MidRed9k8.npy")

3. 压缩存储减小体积

如果对文件体积敏感,可在二进制写入基础上加入gzip压缩,压缩后体积大概率比CSV更小,同时保持较快的读写速度:

C++写入代码(需引入zlib库):

#include <zlib.h>

// 先转为连续一维数组
vector<float> flatData;
// ... 填充flatData逻辑同上 ...

// 写入gzip压缩文件
gzFile gzout = gzopen("../output/11_14MidRed9k8.bin.gz", "wb");
if (gzout) {
    int width = 9000, height = 9000, channels = 4;
    gzwrite(gzout, &width, sizeof(width));
    gzwrite(gzout, &height, sizeof(height));
    gzwrite(gzout, &channels, sizeof(channels));
    gzwrite(gzout, flatData.data(), flatData.size() * sizeof(float));
    gzclose(gzout);
}

Python读取代码:

import numpy as np
import gzip

with gzip.open("../output/11_14MidRed9k8.bin.gz", "rb") as f:
    width = np.frombuffer(f.read(4), dtype=np.int32)[0]
    height = np.frombuffer(f.read(4), dtype=np.int32)[0]
    channels = np.frombuffer(f.read(4), dtype=np.int32)[0]
    data = np.frombuffer(f.read(), dtype=np.float32)
    colorMat = data.reshape((height, width, channels))

4. 优化C++数据结构

当前嵌套vector<vector<vector<float>>>的内存碎片化严重,遍历效率低。可改为:

  • 用vector<vector<array<float,4>>>替代,内存连续性更好,访问更快;
  • 直接使用一维数组存储(按行优先排列:height * width * channels),彻底消除嵌套容器的性能开销。

总结

优先推荐使用numpy的.npy格式,兼顾读写速度和Python端的易用性;若对文件体积敏感,搭配gzip压缩;同时优化C++端的数据结构为连续内存,进一步提升写入效率。

内容的提问来源于stack exchange,提问作者trent

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 01:41:29