遍历Pandas DataFrame按同名列拆分多DataFrame的问题排查
问题解决:按Customer列拆分生成对齐列的多DataFrame
主文件示例

现有代码
d = pd.read_excel(path, header = 1,sheet_name='Master File example') dct={} for i in d.filter(like= "Customer"): dct[f'df {i}'] = d.loc[:,i:'TAMBA'] print(dct)
需求与问题
需求
- 从每个
Customer列(含自动命名的Customer.1、Customer.2等)切片至下一个Customer列的前一列,生成多个DataFrame用于对比 - 若后续Customer列缺失初始Customer列的字段,需为其创建对应空列,保证所有DataFrame列结构一致
- 处理需覆盖到最后一个Customer列
当前问题
现有代码仅生成一个Customer对应的DataFrame,无法正确拆分所有Customer区块
解决方案
步骤1:定位所有Customer列的位置
先获取所有Customer列的索引,明确每个区块的起止边界:
import pandas as pd # 读取Excel文件 d = pd.read_excel(path, header=1, sheet_name='Master File example') # 筛选所有含"Customer"的列名及其索引 customer_cols = [col for col in d.columns if 'Customer' in col] customer_indices = [d.columns.get_loc(col) for col in customer_cols]
步骤2:拆分区块并对齐列结构
遍历每个Customer列,按边界切片,再以第一个Customer区块的列作为基准对齐所有DataFrame:
dct = {} # 取第一个Customer区块的列作为基准列结构 base_columns = d.loc[:, customer_cols[0]:d.columns[customer_indices[1]-1]].columns for idx, col_name in enumerate(customer_cols): # 确定当前区块的结束索引 if idx < len(customer_cols) - 1: end_col_idx = customer_indices[idx+1] - 1 else: end_col_idx = len(d.columns) - 1 # 切片当前Customer区块 current_df = d.loc[:, col_name:d.columns[end_col_idx]] # 对齐列:补充缺失的基准列,空值填充None current_df = current_df.reindex(columns=base_columns, fill_value=None) # 存入字典 dct[f'df {col_name}'] = current_df # 验证结果 for df_name, df in dct.items(): print(f"\n--- {df_name} ---") print(df.head())
代码说明
- 通过
customer_indices精准定位每个Customer列的位置,避免直接用列名切片导致的越界或错误截取问题 - 用
reindex统一所有DataFrame的列结构,确保后续对比时列完全匹配 - 最后一个Customer区块直接截取到表格末尾,无需寻找下一个Customer列
内容的提问来源于stack exchange,提问作者deeplearning
相关产品推荐
相关产品推荐

