如何基于Pandas DataFrame的category列条件生成指定值的新列?
Pandas按条件生成新列的实现方法
原始数据结构
先构造你的原始DataFrame:
import pandas as pd import numpy as np df = pd.DataFrame({ 'category': ['A.', 'B.', 'A.', 'B.', 'C.', 'D.'], 'value1': [20., 40., 60., 80., 10., 20.], 'value2': [30., 50., 70., 90., 10., 20.] })
方法1:嵌套numpy.where
通过嵌套np.where处理双重条件,逻辑简洁:
df['new_col'] = np.where(df['category'] == 'A.', df['value1'], np.where(df['category'] == 'B.', df['value2'], np.nan))
方法2:使用df.loc赋值
先初始化新列为nan,再按条件逐个赋值,直观易懂:
df['new_col'] = np.nan df.loc[df['category'] == 'A.', 'new_col'] = df['value1'] df.loc[df['category'] == 'B.', 'new_col'] = df['value2']
方法3:使用np.select
适合多条件场景,扩展性强,这里同样适用:
conditions = [ df['category'] == 'A.', df['category'] == 'B.' ] choices = [ df['value1'], df['value2'] ] df['new_col'] = np.select(conditions, choices, default=np.nan)
最终输出
执行上述任意方法后,得到的结果如下:
category value1 value2 new_col 0 A. 20. 30. 20. 1 B. 40. 50. 50. 2 A. 60. 70. 60. 3 B. 80. 90. 90. 4 C. 10. 10. NaN 5 D. 20. 20. NaN
内容的提问来源于stack exchange,提问作者Jesse Jin
相关产品推荐
相关产品推荐

