You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++加载pandas导出的多索引列CSV 实现年份为索引指标为多层列咨询

C++加载pandas多索引列CSV实现方案

首先确认你从pandas导出CSV时保留了多索引层级,导出时建议使用以下代码保证格式统一:
df.to_csv("finance_data.csv", index=True, header=True, encoding="utf-8-sig")
导出后的CSV结构默认如下:

,code1,code1,code1,code2,code2,code2...
,营收,净利润,负债率,营收,净利润,负债率...
2014,1234,567,0.34,3245,782,0.29...
2015,...
共8行数据对应8个年份,列前两行分别为code、财务指标两个层级,第一列为年份行索引。


实现步骤

1. 定义数据存储结构

为适配多层级列+行索引的查询需求,定义三类存储变量:

  • 多层列索引映射:std::map<std::pair<std::string, std::string>, int>,键为(code, 财务指标)组合,值为对应数据的列序号
  • 行索引映射:std::map<std::string, int>,键为年份字符串,值为对应数据的行序号
  • 数据矩阵:std::vector<std::vector<double>>存储所有CSV数值内容

2. 核心解析代码

不需要引入重型依赖,原生C++即可实现解析,代码如下:

#include <iostream>
#include <fstream>
#include <vector>
#include <map>
#include <sstream>
#include <string>
#include <utility>

// 按分隔符分割字符串
std::vector<std::string> split(const std::string& s, char delimiter) {
    std::vector<std::string> tokens;
    std::string token;
    std::istringstream tokenStream(s);
    while (std::getline(tokenStream, token, delimiter)) {
        tokens.push_back(token);
    }
    return tokens;
}

int main() {
    std::ifstream file("finance_data.csv");
    std::string line;
    std::map<std::pair<std::string, std::string>, int> col_index;
    std::map<std::string, int> row_index;
    std::vector<std::vector<double>> finance_data;

    // 读取第一层级列头:code
    std::getline(file, line);
    std::vector<std::string> code_level = split(line, ',');
    // 读取第二层级列头:财务指标
    std::getline(file, line);
    std::vector<std::string> indicator_level = split(line, ',');

    // 构建多层列索引,跳过第一列的空表头(对应年份行索引列)
    for (int i = 1; i < code_level.size(); ++i) {
        std::pair<std::string, std::string> col_key = {code_level[i], indicator_level[i]};
        col_index[col_key] = i - 1;
    }

    // 读取数据行,构建行索引与数据矩阵
    int row_cnt = 0;
    while (std::getline(file, line)) {
        std::vector<std::string> row_data = split(line, ',');
        std::string year = row_data[0];
        row_index[year] = row_cnt;
        // 转换数值存入矩阵
        std::vector<double> data_row;
        for (int i = 1; i < row_data.size(); ++i) {
            // 空值可自行替换为默认值如NAN
            data_row.push_back(row_data[i].empty() ? NAN : std::stod(row_data[i]));
        }
        finance_data.push_back(data_row);
        row_cnt++;
    }

    // 测试查询:2015年code为000001的净利润
    double target_val = finance_data[row_index["2015"]][col_index[{"000001", "净利润"}]];
    std::cout << "查询结果:" << target_val << std::endl;

    return 0;
}

3. 注意事项

  • 如果CSV存在特殊字符、转义逗号等复杂格式,建议替换手动分割逻辑为第三方轻量CSV解析库如csv-parser,提升容错性
  • 如有量化计算需求,可将finance_data替换为Eigen矩阵类型,方便后续做矩阵运算、统计分析
  • 行索引如果需要用数值类型,可把row_index的键改为int类型,转换年份字符串为整数即可

内容的提问来源于stack exchange,提问作者hyemin ju

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 16:36:00