如何基于条件映射字典为pandas DataFrame新增匹配结果列
实现方法
首先调整原有映射字典的结构为按匹配优先级排序的规则列表,再通过自定义函数逐行匹配即可得到目标结果,完整代码如下:
import pandas as pd import numpy as np # 原始映射字典 dict_map = { "Yes" : {'Sex':'F', 'Pregnant': 'Y'}, "No" : {'Sex':'F', 'Pregnant': 'N'}, "N/A" : {'Sex': 'M'}, } # 原始DataFrame df = pd.DataFrame( { 'ID': ['AB01', 'AB02', 'AB03', 'AB04', 'AB05','AB06'], 'Sex': ["F","M","M",np.nan,"F","F"], 'Pregnant': ['Y', 'N',np.nan, 'Y', '', 'N'] } ) cols_of_interest = ["Sex","Pregnant"] # 调整为按优先级排序的匹配规则(条件越多优先级越高) match_rules = [(res, cond) for res, cond in dict_map.items()] # 如需自定义匹配顺序可手动定义规则: # match_rules = [("Yes", {'Sex':'F', 'Pregnant': 'Y'}), ("No", {'Sex':'F', 'Pregnant': 'N'}), ("N/A", {'Sex': 'M'})] def match_row(row): for result, conditions in match_rules: is_match = True for col, target in conditions.items(): row_val = row[col] # 过滤空值、空字符串等无效值 if pd.isna(row_val) or str(row_val).strip() == "": is_match = False break if row_val != target: is_match = False break if is_match: return result return np.nan # 生成结果列 df["Result"] = df.apply(match_row, axis=1) # 查看最终输出 print(df)
运行后生成的Result列即为需求的计算结果。
内容的提问来源于stack exchange,提问作者Fazli
相关产品推荐
相关产品推荐

