如何不使用循环为pandas DataFrame多列同时应用各自的映射规则
解决方案
你可以直接使用pandas原生的DataFrame.replace()方法实现需求,该方法原生支持传入列名到映射规则的嵌套字典,无需自行编写遍历逻辑:
import pandas as pd # 原始数据 df = pd.DataFrame({'col1': ['a','a','b','b'], 'col2': ['a','a','b','b']}) dict_map = {'col1': {'a': 0.5, 'b': 1.0}, 'col2': {'a': 0.6, 'b': 0.9}} # 核心替换逻辑 df = df.replace(dict_map)
运行后输出结果如下:
col1 col2 0 0.5 0.6 1 0.5 0.6 2 1.0 0.9 3 1.0 0.9
补充说明
如果数据量较大,想要更高的运行效率,可以使用map结合列名匹配的方式实现,性能优于replace:
df = df.apply(lambda col: col.map(dict_map[col.name]))
注意如果列中存在映射字典中没有的取值,replace会保留原始值,map会返回NaN,可根据实际需求选择。
内容的提问来源于stack exchange,提问作者Aaron England
相关产品推荐
相关产品推荐

