如何在Python中实现两个数据集间的数值映射?
数值映射实现方案
步骤说明
- 先从df2中为每个变量(A1、A2)构建数值到文本的映射字典
- 遍历df1中需要替换的列,使用对应字典完成值的替换
完整代码示例
import pandas as pd # 构建示例df1 df1 = pd.DataFrame({ 'Pers.': [500, 501], 'A1': [1, 3], 'A2': [2, 1] }) # 构建示例df2 df2 = pd.DataFrame({ 'VAR': ['A1', 'A1', 'A1', 'A2', 'A2'], 'RES': [1, 2, 3, 1, 2], 'MEAN': ['Yes', 'No', 'Maybe', 'answered', 'not answered'] }) # 生成映射字典:key是变量名(如A1),value是{RES: MEAN}的字典 mapping = df2.groupby('VAR').apply(lambda x: dict(zip(x['RES'], x['MEAN']))).to_dict() # 遍历df1的列,完成替换 for col in df1.columns: if col in mapping: df1[col] = df1[col].replace(mapping[col]) print(df1)
运行结果
| Pers. | A1 | A2 |
|---|---|---|
| 500 | Yes | not answered |
| 501 | Maybe | answered |
关键逻辑解释
groupby('VAR').apply(...):按变量名分组,为每个组生成RES到MEAN的映射关系replace(mapping[col]):批量替换列中的所有匹配值,比逐单元格遍历效率更高
内容的提问来源于stack exchange,提问作者yasyes
相关产品推荐
相关产品推荐

