如何叠加两个DataFrame并为历史差异值添加前缀?
问题分析与解决方法
错误原因
你的代码里更新字典的逻辑有误:dict_now.update({dict_now[mother_key][son_key]: '(history) ' + value}) 这行是把当前dict_now[mother_key][son_key]的值(也就是nan)作为外层字典的键添加新条目,而非修改内层字典中son_key对应的值。这就是为什么结果里多出一堆nan: '(history) ...'的无效条目,目标内层字典的空值却没被替换。
方法一:修正字典操作逻辑
直接修改内层字典的对应键值,替代错误的外层update操作:
import pandas as pd from io import StringIO import math csvfile_now = StringIO( """Name Project_A Project_B Project_C Project_D Mike 2 8 Jane 7 Kate 17 """) csvfile_history = StringIO( """Name Project_A Project_B Project_C Project_D Mike 7 Jane 8 2 6 Kate 11 12 1 """) df_now = pd.read_csv(csvfile_now, sep='\t', engine='python') df_history = pd.read_csv(csvfile_history, sep='\t', engine='python') df_now = df_now.set_index('Name') df_history = df_history.set_index('Name') dict_now = df_now.to_dict('index') dict_history = df_history.to_dict('index') # 遍历并更新字典 for name, projects in dict_now.items(): for proj, val in projects.items(): if math.isnan(val) and not math.isnan(dict_history[name][proj]): # 直接修改内层字典的目标项目值 dict_now[name][proj] = f'(history) {dict_history[name][proj]}' # 转换回DataFrame result_df = pd.DataFrame.from_dict(dict_now, orient='index').reset_index().rename(columns={'index': 'Name'}) print(result_df)
运行输出:
Name Project_A Project_B Project_C Project_D 0 Mike 2.0 (history) 7.0 NaN 8.0 1 Jane NaN 7.0 (history) 2.0 (history) 6.0 2 Kate 17.0 NaN (history) 12.0 (history) 1.0
方法二:Pandas原生向量化操作(推荐)
无需转字典,直接利用Pandas的批量处理能力,代码更简洁高效:
import pandas as pd from io import StringIO csvfile_now = StringIO( """Name Project_A Project_B Project_C Project_D Mike 2 8 Jane 7 Kate 17 """) csvfile_history = StringIO( """Name Project_A Project_B Project_C Project_D Mike 7 Jane 8 2 6 Kate 11 12 1 """) df_now = pd.read_csv(csvfile_now, sep='\t', engine='python').set_index('Name') df_history = pd.read_csv(csvfile_history, sep='\t', engine='python').set_index('Name') # 定位需要填充的单元格:df_now为空且df_history有值 mask = df_now.isna() & df_history.notna() # 批量替换目标单元格,其余保留原df_now的值 result_df = df_now.mask(mask, '(history) ' + df_history.astype(str)).reset_index() print(result_df)
该方法通过mask函数精准定位待替换区域,实现批量处理,完全符合Pandas的最佳实践,输出结果与方法一一致。
内容的提问来源于stack exchange,提问作者Mark K
相关产品推荐
相关产品推荐

