Pandas DataFrame按指定条件通过字典替换列值的实现问题
原有代码问题说明
你写的代码有三个核心错误:
df.set_index('name')执行后会返回新的DataFrame,你没有赋值回原变量,原df的索引还是默认的0-4整数,循环中用累加的row取值会直接索引越界- 循环逻辑只处理了
ha列,没有覆盖dict1中所有需要处理的列,且取dict1[i]时拿到的是对应列的整行字典,不是当前行的匹配值 - 手动遍历行的操作效率极低,Pandas原生向量化操作完全可以满足需求,不需要手动循环
正确实现代码
import pandas as pd # 构造原始DataFrame,修正索引设置 data = {'name':['sam','rye','lori','chris','sara'], 'ha':[0.020,1,0.05,0.7,0.001], 'he':[1,1,0.1,0.0001,1], 'hi':[0.001,0.002,0.0021,0.3,0.005], 'ho':[0.0002,0.0043,0.0067,0.0123,0.0110], 'hu':[0.7500,0.0540,0.0030,1,0.0081], 'hm':[0.002,0.0021,0.3,0.005,1]} df = pd.DataFrame(data) # 索引设置必须赋值回原变量才生效 df = df.set_index('name') # 构造替换规则字典 dict1 = {'ha': { 'sam' : 0.020, 'rye' : -0.018, 'lori': 0.05, 'chris': 0.7, 'sara' : 0.001}, 'he': { 'sam' : 0.00005, 'rye' : 0, 'lori': 1, 'chris': -2, 'jesse' : 5}} # 将替换规则转为DataFrame,自动对齐行、列 df_replace = pd.DataFrame(dict1) # 取公共行列,仅处理匹配到的行和列,不存在的行/列自动跳过 common_cols = df.columns.intersection(df_replace.columns) common_index = df.index.intersection(df_replace.index) # 按规则替换:原值大于替换值时用替换值,否则保留原值 df.loc[common_index, common_cols] = df.loc[common_index, common_cols].where( df.loc[common_index, common_cols] <= df_replace.loc[common_index, common_cols], df_replace.loc[common_index, common_cols] )
替换后结果
ha he hi ho hu hm name sam 0.020 0.00005 0.0010 0.0002 0.7500 0.0020 rye -0.018 0.00000 0.0020 0.0043 0.0540 0.0021 lori 0.050 0.10000 0.0021 0.0067 0.0030 0.3000 chris 0.700 -2.00000 0.3000 0.0123 1.0000 0.0050 sara 0.001 1.00000 0.0050 0.0110 0.0081 1.0000
内容的提问来源于stack exchange,提问作者newbzzs
相关产品推荐
相关产品推荐

