如何在Pandas DataFrame中根据条件为指定列赋值?
首先修正你提供的代码中的语法错误(字典项缺少逗号、未导入numpy),修正后的代码如下:
import pandas as pd import numpy as np df = pd.DataFrame({ 'action': ['visited', 'clicked', 'switched'], 'target': ['pricing page', 'homepage', 'succeesed'], 'type': [np.nan, np.nan, np.nan] })
针对你需要根据条件为type列赋值的需求,这里提供几种实用方法:
方法1:使用loc直接赋值(推荐,高效直观)
通过布尔索引定位符合条件的行,直接修改type列的值:
df.loc[(df['action'] == 'visited') & (df['target'] == 'pricing page'), 'type'] = 'free'
方法2:使用np.where设置条件赋值
如果需要同时处理不符合条件的情况(比如保留原值),可以用numpy的where函数:
df['type'] = np.where( (df['action'] == 'visited') & (df['target'] == 'pricing page'), 'free', df['type'] # 不符合条件时保留原值 )
方法3:使用apply处理复杂逻辑
如果后续有更复杂的多条件判断,可使用apply逐行处理(注意:大数据集下效率低于前两种方法):
def set_type(row): if row['action'] == 'visited' and row['target'] == 'pricing page': return 'free' return row['type'] df['type'] = df.apply(set_type, axis=1)
执行上述任意方法后,目标行的type列会被设置为free,其余行保留原空值。
内容的提问来源于stack exchange,提问作者bengisu yilmaz
相关产品推荐
相关产品推荐

