You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历Pandas DataFrame按同名列拆分多DataFrame的问题排查

问题解决:按Customer列拆分生成对齐列的多DataFrame

主文件示例

Master File Example

现有代码

d = pd.read_excel(path, header = 1,sheet_name='Master File example')
dct={}
for i in d.filter(like= "Customer"):
   dct[f'df {i}'] = d.loc[:,i:'TAMBA']
print(dct)

需求与问题

需求

  • 从每个Customer列(含自动命名的Customer.1、Customer.2等)切片至下一个Customer列的前一列,生成多个DataFrame用于对比
  • 若后续Customer列缺失初始Customer列的字段,需为其创建对应空列,保证所有DataFrame列结构一致
  • 处理需覆盖到最后一个Customer列

当前问题

现有代码仅生成一个Customer对应的DataFrame,无法正确拆分所有Customer区块


解决方案

步骤1:定位所有Customer列的位置

先获取所有Customer列的索引,明确每个区块的起止边界:

import pandas as pd

# 读取Excel文件
d = pd.read_excel(path, header=1, sheet_name='Master File example')

# 筛选所有含"Customer"的列名及其索引
customer_cols = [col for col in d.columns if 'Customer' in col]
customer_indices = [d.columns.get_loc(col) for col in customer_cols]

步骤2:拆分区块并对齐列结构

遍历每个Customer列,按边界切片,再以第一个Customer区块的列作为基准对齐所有DataFrame:

dct = {}
# 取第一个Customer区块的列作为基准列结构
base_columns = d.loc[:, customer_cols[0]:d.columns[customer_indices[1]-1]].columns

for idx, col_name in enumerate(customer_cols):
    # 确定当前区块的结束索引
    if idx < len(customer_cols) - 1:
        end_col_idx = customer_indices[idx+1] - 1
    else:
        end_col_idx = len(d.columns) - 1
    
    # 切片当前Customer区块
    current_df = d.loc[:, col_name:d.columns[end_col_idx]]
    # 对齐列:补充缺失的基准列,空值填充None
    current_df = current_df.reindex(columns=base_columns, fill_value=None)
    # 存入字典
    dct[f'df {col_name}'] = current_df

# 验证结果
for df_name, df in dct.items():
    print(f"\n--- {df_name} ---")
    print(df.head())

代码说明

  • 通过customer_indices精准定位每个Customer列的位置,避免直接用列名切片导致的越界或错误截取问题
  • 用reindex统一所有DataFrame的列结构,确保后续对比时列完全匹配
  • 最后一个Customer区块直接截取到表格末尾,无需寻找下一个Customer列

内容的提问来源于stack exchange,提问作者deeplearning

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 07:24:18