如何识别两个DataFrame的唯一元素并追加对应新行
解决方案
你可以直接用pd.concat拼接新行,新行仅指定第一列取值时,其余列会自动填充NaN,完整可运行的函数如下:
import pandas as pd import numpy as np def standardize_shape(df1, df2): # 取第一列的列名,适配任意列名的场景 first_col = df1.columns[0] # 分别提取两个DataFrame第一列的唯一值 set1 = set(df1.iloc[:, 0]) set2 = set(df2.iloc[:, 0]) # 计算仅在各自表中存在的元素 only_df1 = set1 - set2 only_df2 = set2 - set1 # 批量给df1补全缺失元素 if only_df2: append_rows = pd.DataFrame({first_col: list(only_df2)}) df1 = pd.concat([df1, append_rows], ignore_index=True) # 批量给df2补全缺失元素 if only_df1: append_rows = pd.DataFrame({first_col: list(only_df1)}) df2 = pd.concat([df2, append_rows], ignore_index=True) return df1, df2
测试验证
# 构造你示例中的测试数据 d1 = {'col1': [1, 2, 5], 'col2': [3, 4, 6]} df1 = pd.DataFrame(data=d1) d2 = {'col1': [1, 2, 6], 'col2': [3, 4, 7]} df2 = pd.DataFrame(data=d2) df1_res, df2_res = standardize_shape(df1, df2) print(df1_res) print("-----") print(df2_res)
输出结果和你预期完全一致:
col1 col2 0 1 3.0 1 2 4.0 2 5 6.0 3 6 NaN ----- col1 col2 0 1 3.0 1 2 4.0 2 6 7.0 3 5 NaN
内容的提问来源于stack exchange,提问作者Jengels
相关产品推荐
相关产品推荐

