You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

std::unique_ptr使用疑问:未修改指向对象及生存期问题

多线程加载JSON数据的指针使用问题

问题背景

我在项目中通过多线程从JSON文件加载大量数据:一个线程将文件内容加载到std::deque,另一个线程从该deque格式化数据。但使用std::unique_ptr访问deque时,对unique_ptr的修改并未同步到原deque。目前已用引用解决加载问题,但因后续需要随机打乱pic_val_pairs,无法继续使用引用。

核心疑问

  • 使用std::make_unique初始化的std::unique_ptr(或std::shared_ptr)是否会修改初始化时的原对象?
  • std::unique_ptr是否像裸指针一样依赖指向的对象?将指向加载数据的unique_ptr传给其他函数时,是否需要存储原数据,还是仅存储unique_ptr即可(即使原数据被销毁)?
  • 使用unique_ptr后内存占用远高于JSON文件总大小,怀疑它复制了对象,这个判断是否正确?

相关代码

#include <mutex>
#include <thread>
#include <deque>
#include <array>
#include <fstream>

typedef std::deque<std::pair<array<array<double, 28>, 28>, unsigned int>> threadLoadList;
typedef vector<std::pair < std::shared_ptr<vector<double>>, std::shared_ptr<vector<double>> >> dataPointer;


namespace data {
    std::mutex mtx{};
    vector<vector<double>> training_desired_output{};
    vector<vector<double>> training_images{};
    dataPointer pic_val_pairs{};
}

void loadTrainImages(std::unique_ptr<threadLoadList> storage) {
    for(int i = 0; i < 60000; ++i) {
        std::ifstream str(data::training_data_path + "\\image" + Format_number(i) + ".json");
        nlohmann::json data = nlohmann::json::parse(str);
        data::mtx.lock();
        storage->push_back(std::make_pair<array<array<double, 28>, 28>, unsigned int>(
            data["pic"],
            data["val"]
        ));
        data::mtx.unlock();
        str.close();
    }
}

int main() {
    threadLoadList pics_vals{};

    auto load = std::thread(loadTrainImages, std::make_unique<threadLoadList>(pics_vals));

    while(true) {
        data::mtx.lock();
        unsigned int size = pics_vals.size();
        data::mtx.unlock();
        if(size > 0) {
            std::cout << "got one\n";
            data::mtx.lock();
            std::pair<array<array<double, 28>, 28>, unsigned int> copy = pics_vals.front();
            pics_vals.pop_front();
            data::mtx.unlock();

            //format 28 x 28 array
            vector<double> pic{};
            for(int i = 27; i > -1; --i) {
                for(int n = 0; n < 28; ++n) {
                    pic.push_back(copy.first[i][n]);
                }
            }

            // format val-array
            vector<double> val(10);
            val = { 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0 };
            val[copy.second] = 1.0;

            // do I need to store the data (pic and val) pic_val_pairs is pointing to globally? Or is it enought, to store the pointers to them and have pic and val being destroyed?
            data::training_images.push_back(pic);
            data::training_desired_output.push_back(val);
            data::pic_val_pairs.push_back(
                std::make_pair<std::shared_ptr<vector<double>>, std::shared_ptr<vector<double>>>(
                std::make_shared<vector<double>>(data::training_images.back()),
                std::make_shared<vector<double>>(data::training_desired_output.back())
            ));
        }
        if(data::pic_val_pairs.size() >= 60000) {
            break;
        }
    }
    // Do something
}

问题解答

1. std::make_unique初始化的指针是否会修改原对象?

不会。你代码中std::make_unique<threadLoadList>(pics_vals)是复制了pics_vals这个deque的完整副本,unique_ptr指向的是这个独立副本而非原deque。因此加载线程往副本中添加数据时,原pics_vals完全不会受到影响——这就是你看不到数据同步的根本原因。

2. std::unique_ptr是否依赖指向的对象?

是的。unique_ptr本质是带生命周期管理的裸指针,它仅负责自动释放指向的对象,本身不存储对象数据。如果原对象被销毁,unique_ptr会变成悬空指针,访问它会触发未定义行为。

传递unique_ptr给其他函数时:

  • 如果unique_ptr指向的是通过make_unique新建的对象,只要unique_ptr未被销毁,对象就会保持存在,无需额外存储原数据;
  • 如果unique_ptr指向外部对象(比如全局变量),必须保证外部对象在unique_ptr的使用周期内始终存在,否则会引发错误。

3. 内存占用过高的原因

你的判断正确,确实存在不必要的复制操作:

  • 初始化unique_ptr时复制了整个deque;
  • std::make_shared<vector<double>>(data::training_images.back())又复制了training_images中的vector,等于一份数据同时存储在全局容器和shared_ptr指向的内存中,双倍占用内存。

代码修复建议

解决加载线程与主线程的同步问题

不要传递deque的副本,让加载线程直接操作原deque,可以通过引用传递实现:

// 修改函数参数为引用
void loadTrainImages(threadLoadList& storage) {
    for(int i = 0; i < 60000; ++i) {
        std::ifstream str(data::training_data_path + "\\image" + Format_number(i) + ".json");
        nlohmann::json data = nlohmann::json::parse(str);
        std::lock_guard<std::mutex> lock(data::mtx); // 用lock_guard自动管理锁,避免遗漏解锁
        storage.push_back(std::make_pair<array<array<double, 28>, 28>, unsigned int>(
            data["pic"],
            data["val"]
        ));
        str.close();
    }
}

// main中创建线程时传递原deque的引用
auto load = std::thread(loadTrainImages, std::ref(pics_vals));

解决内存占用过高与随机打乱需求

如果要打乱pic_val_pairs,可以直接存储全局容器的指针或索引,避免数据复制:

// 修改dataPointer为存储指针类型
typedef vector<std::pair<vector<double>*, vector<double>*>> dataPointer;

// 插入时直接取全局容器元素的地址
data::pic_val_pairs.push_back(
    std::make_pair(&data::training_images.back(), &data::training_desired_output.back())
);

// 打乱操作直接作用于pic_val_pairs,不影响原数据存储
std::shuffle(data::pic_val_pairs.begin(), data::pic_val_pairs.end(), std::mt19937(std::random_device{}()));

若一定要使用shared_ptr,可以通过移动语义避免复制:

// 将pic和val移动进全局容器,避免复制
data::training_images.emplace_back(std::move(pic));
data::training_desired_output.emplace_back(std::move(val));
// 自定义删除器,让shared_ptr不释放全局容器的内存(由全局容器管理生命周期)
data::pic_val_pairs.push_back(
    std::make_pair(
        std::shared_ptr<vector<double>>(&data::training_images.back(), [](auto*){}),
        std::shared_ptr<vector<double>>(&data::training_desired_output.back(), [](auto*){})
    )
);

内容的提问来源于stack exchange,提问作者user27097105

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 17:45:04