Pandas DataFrame公式计算时NaN值的更优处理方法是什么
解决方案
你不需要修改indicator_df,只需要在计算前判断表达式是否为空,为空时直接填充默认值即可,具体修改逻辑如下:
- 遍历
input_df所有行,而非仅第一行 - 读取每个
KPI_Key的表达式后先判断是否为NaN,为空时直接给对应列赋值0(和你给出的样例输出匹配,也可根据需求换成pd.NA等其他默认值) - 非空表达式再调用
eval执行计算 - 替换已废弃的
append方法,用列表拼接+pd.concat提升性能
import pandas as pd import numpy as np import re # 原始数据构造 input_df = pd.DataFrame({'KPI_ID': ['A','B','C','D','E'], 'KPI_Key1': ['(C602+C603)','(C605+C606)','75','(32*(C603+44))','L239'] , 'KPI_Key2' : ['C601','C602','L239+C602','75',np.NaN] , 'KPI_Key3' : ['75',np.NaN,np.NaN,np.NaN,'C601']}) indicator_df = pd.DataFrame({'PatientID': [1,2,3,4,5], '99' : ['1','0','1','0','1'], '75' : ['0','0','1','0','0'], 'C604' : ['1','0','1','0','1'], 'C602' : ['0','0','1','0','1'], 'C601' : ['1','0','0','0','1'], 'C603' : ['0','0','1','1','1'], 'C605' : ['0','1','1','0','0'], 'C606' : ['0','1','1','1','1'], '44' : ['1','0','1','0','1'], 'L239' : ['0','0','1','1','1'], '32' : ['1','0','1','0','1'], }).set_index('PatientID').astype('int32') # 核心计算逻辑 out_list = [] key_cols = ['KPI_Key1','KPI_Key2','KPI_Key3'] for _, row in input_df.iterrows(): kpi_id = row['KPI_ID'] # 初始化当前KPI的结果表 res = indicator_df.reset_index()[['PatientID']].copy() res['KPI_ID'] = kpi_id for col in key_cols: exp = row[col] if pd.isna(exp): # 空表达式直接填默认值,可按需修改 res[col] = 0 else: # 格式化表达式后执行计算 formatted_exp = re.sub(r'(\w+)', r'`\1`', exp) res[col] = indicator_df.eval(formatted_exp).values out_list.append(res) # 合并所有结果得到最终输出 final_out_df = pd.concat(out_list, ignore_index=True)
运行上述代码后输出的final_out_df与要求的格式完全匹配,空表达式对应的列不会触发执行错误。
内容的提问来源于stack exchange,提问作者user3234112
相关产品推荐
相关产品推荐

