基于列名与位置为pandas DataFrame创建多级列索引的实现方法
pandas多级列索引实现方案
直接按以下步骤实现即可,无需硬编码中间列名称,完全匹配你的位置指定需求:
完整实现代码
import pandas as pd # 1. 定位maths类列的起始位置 cols = df.columns.tolist() maths_start_idx = next(i for i, col in enumerate(cols) if col.startswith("maths_")) # 2. 构造多级列索引元组列表 multi_columns = [] # 处理id列(保持独立) multi_columns.append(("", "id")) # 处理student_data分类下的列:第二列到maths前一列 for col in cols[1:maths_start_idx]: multi_columns.append(("student_data", col)) # 处理学科成绩列,拆分0级学科、1级指标 for col in cols[maths_start_idx:]: subject, metric = col.split("_", maxsplit=1) multi_columns.append((subject, metric)) # 3. 替换原df的列索引 df.columns = pd.MultiIndex.from_tuples(multi_columns)
方案说明
- 自动定位
maths相关列的起始位置,中间name到class不管有多少列都可以自动归入student_data分类,无需手动指定列名 - 成绩列只拆分第一个下划线,即使后续指标名包含下划线也不会出错
- id列0级索引设为空字符串,完全符合你给出的最终结构展示样式,如果需要id完全保持单级样式,也可以将id列设为行索引:
df = df.set_index("id"),后续列的多级索引结构不变
内容的提问来源于stack exchange,提问作者Ankur
相关产品推荐
相关产品推荐

