如何基于条件使用嵌套字典为Pandas Dataframe映射新列值
实现方案
核心思路
按dict_map的键顺序逐行匹配规则,命中首个符合条件的规则后立即返回对应键值,无需后续校验。Python 3.7及以上版本原生支持字典插入顺序保留,低版本可将规则调整为列表嵌套元组的形式保证顺序。
完整代码
import pandas as pd import numpy as np # 初始化原始数据 dict_map = { 'Anti' : {'Drug':('A','B','C')}, 'Undef': {'Drug':'D','Name':'Type X'}, 'Vit ' : {'Name': 'Vitamin C'}, 'Placebo Effect' : {'Name':'Placebo', 'Batch':'XYZ'}, } df = pd.DataFrame( { 'ID': ['AB01', 'AB02', 'AB03', 'AB04', 'AB05','AB06'], 'Drug': ["A","B","A",np.nan,"D","D"], 'Name': ['Placebo', 'Vitamin C', np.nan, 'Placebo', '', 'Type X'], 'Batch' : ['ABC',np.nan,np.nan,'XYZ',np.nan,np.nan], } ) # 单条规则匹配校验 def match_row(row, rule): for col, cond in rule.items(): row_val = row[col] # 规则为元组时判断值是否属于元组 if isinstance(cond, tuple): if row_val not in cond: return False # 规则为普通值时判断全等 else: if row_val != cond: return False return True # 逐行遍历规则获取结果 def get_result(row): for res_name, rule in dict_map.items(): if match_row(row, rule): return res_name # 无匹配时返回空值 return np.nan # 新增Result列 df['Result'] = df.apply(get_result, axis=1) print(df)
输出结果
| ID | Drug | Name | Batch | Result |
|---|---|---|---|---|
| AB01 | A | Placebo | ABC | Anti |
| AB02 | B | Vitamin C | NaN | Anti |
| AB03 | A | NaN | NaN | Anti |
| AB04 | NaN | Placebo | XYZ | Placebo Effect |
| AB05 | D | NaN | NaN | |
| AB06 | D | Type X | NaN | Undef |
注意事项
- 如果规则中需要匹配空值
np.nan,需单独处理判断逻辑(因为np.nan != np.nan),可新增判断分支:if pd.isna(row_val) and pd.isna(cond): continue - Python 3.7以下版本需将
dict_map替换为collections.OrderedDict或列表格式[('Anti', {'Drug':('A','B','C')}), ...]保证规则优先级顺序 - 数据量超过10万行时可提前将规则向量化处理提升性能,避免逐行
apply的开销
内容的提问来源于stack exchange,提问作者Fazli
相关产品推荐
相关产品推荐

