C++加载pandas导出的多索引列CSV 实现年份为索引指标为多层列咨询
C++加载pandas多索引列CSV实现方案
首先确认你从pandas导出CSV时保留了多索引层级,导出时建议使用以下代码保证格式统一:df.to_csv("finance_data.csv", index=True, header=True, encoding="utf-8-sig")
导出后的CSV结构默认如下:
,code1,code1,code1,code2,code2,code2...
,营收,净利润,负债率,营收,净利润,负债率...
2014,1234,567,0.34,3245,782,0.29...
2015,...
共8行数据对应8个年份,列前两行分别为code、财务指标两个层级,第一列为年份行索引。
实现步骤
1. 定义数据存储结构
为适配多层级列+行索引的查询需求,定义三类存储变量:
- 多层列索引映射:
std::map<std::pair<std::string, std::string>, int>,键为(code, 财务指标)组合,值为对应数据的列序号 - 行索引映射:
std::map<std::string, int>,键为年份字符串,值为对应数据的行序号 - 数据矩阵:
std::vector<std::vector<double>>存储所有CSV数值内容
2. 核心解析代码
不需要引入重型依赖,原生C++即可实现解析,代码如下:
#include <iostream> #include <fstream> #include <vector> #include <map> #include <sstream> #include <string> #include <utility> // 按分隔符分割字符串 std::vector<std::string> split(const std::string& s, char delimiter) { std::vector<std::string> tokens; std::string token; std::istringstream tokenStream(s); while (std::getline(tokenStream, token, delimiter)) { tokens.push_back(token); } return tokens; } int main() { std::ifstream file("finance_data.csv"); std::string line; std::map<std::pair<std::string, std::string>, int> col_index; std::map<std::string, int> row_index; std::vector<std::vector<double>> finance_data; // 读取第一层级列头:code std::getline(file, line); std::vector<std::string> code_level = split(line, ','); // 读取第二层级列头:财务指标 std::getline(file, line); std::vector<std::string> indicator_level = split(line, ','); // 构建多层列索引,跳过第一列的空表头(对应年份行索引列) for (int i = 1; i < code_level.size(); ++i) { std::pair<std::string, std::string> col_key = {code_level[i], indicator_level[i]}; col_index[col_key] = i - 1; } // 读取数据行,构建行索引与数据矩阵 int row_cnt = 0; while (std::getline(file, line)) { std::vector<std::string> row_data = split(line, ','); std::string year = row_data[0]; row_index[year] = row_cnt; // 转换数值存入矩阵 std::vector<double> data_row; for (int i = 1; i < row_data.size(); ++i) { // 空值可自行替换为默认值如NAN data_row.push_back(row_data[i].empty() ? NAN : std::stod(row_data[i])); } finance_data.push_back(data_row); row_cnt++; } // 测试查询:2015年code为000001的净利润 double target_val = finance_data[row_index["2015"]][col_index[{"000001", "净利润"}]]; std::cout << "查询结果:" << target_val << std::endl; return 0; }
3. 注意事项
- 如果CSV存在特殊字符、转义逗号等复杂格式,建议替换手动分割逻辑为第三方轻量CSV解析库如
csv-parser,提升容错性 - 如有量化计算需求,可将
finance_data替换为Eigen矩阵类型,方便后续做矩阵运算、统计分析 - 行索引如果需要用数值类型,可把
row_index的键改为int类型,转换年份字符串为整数即可
内容的提问来源于stack exchange,提问作者hyemin ju
相关产品推荐
相关产品推荐

