如何用Pandas df.mask()实现基于外部变量条件的数据掩码?
用Pandas的mask()替代循环实现位置匹配掩码需求
原始数据与需求
数据定义
import pandas as pd import numpy as np seller1 = [5, 4, 3] # 卖家1在时间1、2、3的位置 seller2 = [4, 2, 1] # 卖家2在时间1、2、3的位置 df = {'customer': [1, 1, 1, 2, 2, 2], 'time': [1,2,3,1,2,3], 'location': [3,4,2,4,3,3], 'demand':[10,12,15,20,8,16], 'price':[3,4,4,5,2,1]} df = pd.DataFrame(df)
需求说明
当某时间点客户的location与任一卖家在该时间的位置不匹配时,将对应行的demand和price掩码为NaN;若匹配则保留原值。
原实现的问题
原代码用for循环逐行判断,不仅效率低下,还会触发SettingWithCopyWarning:
for i in range(df.shape[0]): if df["location"][i] != seller1[int(df["time"][i])-1] and df["location"][i] != seller2[int(df["time"][i])-1]: df["demand"][i] = np.nan df["price"][i] = np.nan
用df.mask()的高效实现
步骤1:生成各时间点卖家的位置映射
把seller1和seller2转换成与df['time']对应的位置序列,方便后续向量化比较:
# 根据time列获取对应时间点的卖家位置 seller1_loc = df['time'].map(lambda x: seller1[x-1]) seller2_loc = df['time'].map(lambda x: seller2[x-1])
步骤2:构建掩码条件
掩码条件为:客户位置既不等于卖家1的位置,也不等于卖家2的位置:
mask_condition = ~((df['location'] == seller1_loc) | (df['location'] == seller2_loc))
步骤3:用mask()应用掩码
直接对需要处理的列应用mask(),满足条件时替换为NaN:
df[['demand', 'price']] = df[['demand', 'price']].mask(mask_condition)
完整代码
import pandas as pd import numpy as np seller1 = [5, 4, 3] seller2 = [4, 2, 1] df = {'customer': [1, 1, 1, 2, 2, 2], 'time': [1,2,3,1,2,3], 'location': [3,4,2,4,3,3], 'demand':[10,12,15,20,8,16], 'price':[3,4,4,5,2,1]} df = pd.DataFrame(df) # 生成对应时间的卖家位置序列 seller1_loc = df['time'].apply(lambda x: seller1[x-1]) seller2_loc = df['time'].apply(lambda x: seller2[x-1]) # 构建掩码条件 mask = ~((df['location'] == seller1_loc) | (df['location'] == seller2_loc)) # 应用掩码 df[['demand', 'price']] = df[['demand', 'price']].mask(mask) print(df)
输出结果
customer time location demand price 0 1 1 3 NaN NaN 1 1 2 4 12.0 4.0 2 1 3 2 NaN NaN 3 2 1 4 20.0 5.0 4 2 2 3 NaN NaN 5 2 3 3 16.0 1.0
为什么这样更优
- 向量化操作:避免逐行循环,Pandas向量化方法处理大数据集时效率提升显著
- 避免警告:直接对DataFrame列赋值,不会触发
SettingWithCopyWarning - 代码更简洁:逻辑清晰,可读性更强
内容的提问来源于stack exchange,提问作者A Doe
相关产品推荐
相关产品推荐

