You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++读取大数据集二进制文件时抛出std::bad_alloc异常

问题排查与解决方案

核心问题分析

出现std::bad_alloc的核心原因是二进制读写逻辑完全错误,导致数据解析混乱,直接触发内存分配异常:

1. std::string的序列化错误

std::string是动态容器,内部存储的是指向堆内存的指针而非字符串本身。你的write函数直接把Player对象的内存(包括std::string的内部指针)写入文件,读取时直接将文件内容拷贝到Player对象内存中,会导致:

  • name的内部指针指向无效内存,触发未定义行为
  • 后续操作破坏内存结构,间接引发内存分配异常

2. 读写的固定大小不匹配

  • write函数写入了整个Player对象的大小(包含std::unique_ptr<char[]>的内存):sizeof(Player)
  • read函数却只读取sizeof(Player) - sizeof(std::unique_ptr<char[]>)的内容
    这会导致读取的固定部分数据错位,num字段被读入错误的超大值,后续执行std::make_unique<char[]>(num)时尝试分配远超系统可用内存的空间,直接抛出std::bad_alloc。

3. unique_ptr的冗余写入

write函数中写入了std::unique_ptr<char[]>的内存,这部分是进程内的指针值,在读取进程中完全无效,属于冗余且错误的写入。


修复后的代码实现

必须手动序列化/反序列化std::string和动态数组,避免直接拷贝对象内存:

#include <fstream>
#include <vector>
#include <string>
#include <memory>
#include <iostream>

class Player {
public:
    std::string name;  // Length [3, 15]
    int score;
    size_t id;
    size_t num;
    std::unique_ptr<char[]> p;  // Memory allocated on the heap

    // 修正后的写入函数
    void write(std::ostream& os) const {
        // 先写入字符串长度,再写入字符串内容
        size_t name_len = name.size();
        os.write(reinterpret_cast<const char*>(&name_len), sizeof(name_len));
        os.write(name.data(), name_len);

        // 写入固定大小的基础字段
        os.write(reinterpret_cast<const char*>(&score), sizeof(score));
        os.write(reinterpret_cast<const char*>(&id), sizeof(id));
        os.write(reinterpret_cast<const char*>(&num), sizeof(num));

        // 写入动态数组内容
        os.write(p.get(), num);
    }

    // 修正后的读取函数
    bool read(std::istream& is) {
        // 读取字符串长度,再读取内容
        size_t name_len;
        if (!is.read(reinterpret_cast<char*>(&name_len), sizeof(name_len))) {
            return false;
        }
        name.resize(name_len);
        if (!is.read(name.data(), name_len)) {
            return false;
        }

        // 读取固定大小的基础字段
        if (!is.read(reinterpret_cast<char*>(&score), sizeof(score))) {
            return false;
        }
        if (!is.read(reinterpret_cast<char*>(&id), sizeof(id))) {
            return false;
        }
        if (!is.read(reinterpret_cast<char*>(&num), sizeof(num))) {
            return false;
        }

        // 分配并读取动态数组
        p = std::make_unique<char[]>(num);
        if (!is.read(p.get(), num)) {
            return false;
        }
        return true;
    }
};

int main() {
    std::string filename = "file";
    std::ifstream file(filename, std::ios::binary);
    if (!file) {
        std::cerr << "Failed to open file." << std::endl;
        return 1;
    }

    std::vector<Player> players;
    players.reserve(2000000);  // 预分配内存避免多次扩容

    Player player;
    while (player.read(file)) {
        players.push_back(std::move(player));
    }

    std::cout << "Loaded " << players.size() << " players." << std::endl;
    return 0;
}

内存优化建议(针对200万对象场景)

如果数据仅用于排序和导航,可进一步降低内存开销:

  • 替换std::string为固定大小数组:因为name长度固定在3-15之间,用char name[16];替代std::string,避免std::string的额外内存开销(每个std::string至少占24字节,固定数组仅16字节)
  • 使用自定义分配器:给vector搭配内存池分配器(比如boost::pool_allocator),减少内存碎片
  • 内存映射文件:如果内存压力仍大,可以直接将文件映射到内存,按需访问对象,避免一次性加载所有数据到内存

额外注意事项

  • 跨平台兼容性:不同编译器、平台的内存布局(字节序、对齐方式)不同,若需跨平台使用,必须手动处理字节序(比如用htonl/ntohl转换整数类型)
  • 严格错误检查:读取过程中要逐步骤校验流状态,避免因部分读取导致数据错误

内容的提问来源于stack exchange,提问作者App Key

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 12:05:26