如何遍历Pandas compare()结果行生成指定字段变更报告?
解决方案
1. 先筛选目标列
从compare()返回的多级索引DataFrame中,直接提取你需要的字段(比如A、B、C、E),过滤掉不需要的列:
import pandas as pd # 假设updated_records是compare后的结果 target_fields = ['A', 'B', 'C', 'E'] # 多级索引下按第一层列索引筛选指定字段 filtered_df = updated_records[target_fields]
2. 遍历生成变更报告(两种实用方式)
方式一:直接处理多级索引列
无需修改列结构,直接通过多级索引取值,避开iterrows()的层级混乱:
for item_idx in filtered_df.index: row_data = filtered_df.loc[item_idx] change_list = [] for field in target_fields: # 获取字段的旧值(来自df1的self)和新值(来自df2的other) old_val = row_data[(field, 'self')] new_val = row_data[(field, 'other')] # 只保留有实际变更的记录(排除未变更的NaN值) if not pd.isna(old_val) and not pd.isna(new_val): change_list.append(f"{field}: {old_val} -> {new_val}") # 生成指定格式的报告 if change_list: print(f'"For item XXX_{item_idx} those fields where updated:') print(',\n'.join(change_list) + ',') print(' "') print()
方式二:扁平化列索引后使用iterrows()
如果习惯用iterrows(),可以先把多级列索引转成单层,避免层级混淆:
# 扁平化列名,将(A, self)转为A_self,(A, other)转为A_other flat_df = filtered_df.rename(columns=lambda col: f"{col[0]}_{col[1]}") for item_idx, row in flat_df.iterrows(): change_list = [] for field in target_fields: old_col = f"{field}_self" new_col = f"{field}_other" old_val = row[old_col] new_val = row[new_col] if not pd.isna(old_val) and not pd.isna(new_val): change_list.append(f"{field}: {old_val} -> {new_val}") if change_list: print(f'"For item XXX_{item_idx} those fields where updated:') print(',\n'.join(change_list) + ',') print(' "') print()
关键说明
compare()返回的结果中,未变更的字段对应值为NaN,通过pd.isna()过滤即可只保留有实际变更的内容。- 示例中的
XXX_{item_idx}可以替换为你实际的item唯一标识(比如从原DataFrame中获取的业务ID)。
内容的提问来源于stack exchange,提问作者Exc
相关产品推荐
相关产品推荐

