如何将Pandas DataFrame的Prediction列值替换为指定列对应值
问题:替换Pandas DataFrame中Prediction列的值为指定列的对应值
原始DataFrame:
| A | B | C | D | Prediction |
|---|---|---|---|---|
| stipulation | interrelation | jurisdiction | interpretation | D |
| typically | conceivably | tentatively | desperately | C |
| familiar | imaginative | apparent | logical | A |
| plan | explain | study | discard | B |
预期得到的DataFrame:
| A | B | C | D | Prediction |
|---|---|---|---|---|
| stipulation | interrelation | jurisdiction | interpretation | interpretation |
| typically | conceivably | tentatively | desperately | tentatively |
| familiar | imaginative | apparent | logical | familiar |
| plan | explain | study | discard | explain |
解决方案
这里有几种高效的实现方式:
方法1:使用apply按行取值
如果之前用apply没生效,大概率是没指定axis=1(按行处理),正确写法如下:
df['Prediction'] = df.apply(lambda row: row[row['Prediction']], axis=1)
方法2:使用df.lookup(旧版本Pandas适用)
lookup方法可以直接根据行索引和列名匹配取值,代码更简洁:
df['Prediction'] = df.lookup(df.index, df['Prediction'])
注意:Pandas 1.2.0及以上版本中,lookup已被标记为弃用,推荐用其他方法替代。
方法3:numpy索引(大数据集优先)
如果数据量较大,numpy的索引方式性能会比apply更优:
import numpy as np # 获取Prediction列中每个值对应的列索引 col_indices = df.columns.get_indexer(df['Prediction']) # 按行和列索引提取对应值 df['Prediction'] = df.values[np.arange(len(df)), col_indices]
内容的提问来源于stack exchange,提问作者JAMIE WRIGHT
相关产品推荐
相关产品推荐

