如何使用嵌套字典映射替换Pandas DataFrame指定列的值
实现方法
首先需要预处理嵌套映射字典,把ColA中以元组形式定义的多对一映射展开为单个键值对,再调用pandas的map方法完成列值替换,完整实现代码如下:
import pandas as pd import numpy as np # 原有数据定义 dict_map = { 'ColA' : {('A','Distinct (Highest)','B','C'):'Pass'}, 'ColB' : {'Great': 'A', 'Avg': 'Average', 'F (Absent)': 'Fail'} } cols_to_map = ['ColA','ColB'] new_cols = ['ColA-new','ColB-new'] df = pd.DataFrame( { 'ID': ['AB01', 'AB02', 'AB03', 'AB04', 'AB05','AB06'], 'ColA': ["A","B","A",np.nan, "C",np.nan], 'ColB': ['Great', 'Avg', np.nan, np.nan, 'F (Absent)', np.nan] }) # 第一步:预处理映射字典,展开元组类型的键 processed_mapping = {} for col, mapping_rule in dict_map.items(): expanded_rule = {} for key, val in mapping_rule.items(): # 键为元组时拆分每个元素作为独立键,映射到同一个值 if isinstance(key, tuple): for single_key in key: expanded_rule[single_key] = val else: expanded_rule[key] = val processed_mapping[col] = expanded_rule # 第二步:批量生成映射后的新列 for old_col, new_col in zip(cols_to_map, new_cols): df[new_col] = df[old_col].map(processed_mapping[old_col]) # 输出结果查看 print(df)
输出结果示例
| ID | ColA | ColB | ColA-new | ColB-new | |
|---|---|---|---|---|---|
| 0 | AB01 | A | Great | Pass | A |
| 1 | AB02 | B | Avg | Pass | Average |
| 2 | AB03 | A | NaN | Pass | NaN |
| 3 | AB04 | NaN | NaN | NaN | NaN |
| 4 | AB05 | C | F (Absent) | Pass | Fail |
| 5 | AB06 | NaN | NaN | NaN | NaN |
所有原有列会完整保留,新增的两个映射列自动追加到DataFrame末尾,未匹配到映射规则的空值(NaN)会保持不变。
内容的提问来源于stack exchange,提问作者Fazli
相关产品推荐
相关产品推荐

