如何合并两个DataFrame并保留所有值且不生成_x/_y后缀列?
解决DataFrame同名列合并时避免_x/_y后缀的问题
方法一:用combine_first精准匹配填充
这种方法优先保留第一个DataFrame的非NaN值,仅用第二个DataFrame的值填补空缺,适合两表按主键严格匹配的场景:
import pandas as pd import numpy as np # 构建原始数据 df = pd.DataFrame([['person one', 10], ['person two', np.nan]], columns=['Name', 'Example_column']) df2 = pd.DataFrame([['person one', np.nan], ['person two', 'excused']], columns=['Name', 'Example_column']) # 设置Name为索引,合并后恢复索引 df_final = df.set_index('Name').combine_first(df2.set_index('Name')).reset_index() print(df_final)
输出结果:
Name Example_column 0 person one 10.0 1 person two excused
方法二:用concat+groupby批量处理多列
如果有大量同名列需要合并,或者涉及多个DataFrame,这个方法更高效,无需逐个处理列:
# 行拼接两表,按Name分组后取每组第一个非NaN值 df_final = pd.concat([df, df2]).groupby('Name').first().reset_index() print(df_final)
输出结果和方法一完全一致。
方法说明
combine_first:逻辑更精准,明确以第一个表为基准补全数据,适合两表一对一匹配的场景。concat+groupby:通用性更强,支持任意数量的DataFrame合并,自动处理所有同名列的非NaN值合并。
内容的提问来源于stack exchange,提问作者BPhillips
相关产品推荐
相关产品推荐

