You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas df.mask()实现基于外部变量条件的数据掩码?

用Pandas的mask()替代循环实现位置匹配掩码需求

原始数据与需求

数据定义

import pandas as pd
import numpy as np

seller1 = [5, 4, 3]  # 卖家1在时间1、2、3的位置
seller2 = [4, 2, 1]  # 卖家2在时间1、2、3的位置

df = {'customer': [1, 1, 1, 2, 2, 2], 
      'time': [1,2,3,1,2,3], 
      'location': [3,4,2,4,3,3], 
      'demand':[10,12,15,20,8,16], 
      'price':[3,4,4,5,2,1]}
df = pd.DataFrame(df)

需求说明

当某时间点客户的location与任一卖家在该时间的位置不匹配时,将对应行的demand和price掩码为NaN;若匹配则保留原值。

原实现的问题

原代码用for循环逐行判断,不仅效率低下,还会触发SettingWithCopyWarning:

for i in range(df.shape[0]):
    if df["location"][i] != seller1[int(df["time"][i])-1] and df["location"][i] != seller2[int(df["time"][i])-1]:
        df["demand"][i] = np.nan
        df["price"][i] = np.nan

用df.mask()的高效实现

步骤1:生成各时间点卖家的位置映射

把seller1和seller2转换成与df['time']对应的位置序列,方便后续向量化比较:

# 根据time列获取对应时间点的卖家位置
seller1_loc = df['time'].map(lambda x: seller1[x-1])
seller2_loc = df['time'].map(lambda x: seller2[x-1])

步骤2:构建掩码条件

掩码条件为:客户位置既不等于卖家1的位置,也不等于卖家2的位置:

mask_condition = ~((df['location'] == seller1_loc) | (df['location'] == seller2_loc))

步骤3:用mask()应用掩码

直接对需要处理的列应用mask(),满足条件时替换为NaN:

df[['demand', 'price']] = df[['demand', 'price']].mask(mask_condition)

完整代码

import pandas as pd
import numpy as np

seller1 = [5, 4, 3]
seller2 = [4, 2, 1]

df = {'customer': [1, 1, 1, 2, 2, 2], 
      'time': [1,2,3,1,2,3], 
      'location': [3,4,2,4,3,3], 
      'demand':[10,12,15,20,8,16], 
      'price':[3,4,4,5,2,1]}
df = pd.DataFrame(df)

# 生成对应时间的卖家位置序列
seller1_loc = df['time'].apply(lambda x: seller1[x-1])
seller2_loc = df['time'].apply(lambda x: seller2[x-1])

# 构建掩码条件
mask = ~((df['location'] == seller1_loc) | (df['location'] == seller2_loc))

# 应用掩码
df[['demand', 'price']] = df[['demand', 'price']].mask(mask)

print(df)

输出结果

customer  time  location  demand  price
0         1     1         3     NaN    NaN
1         1     2         4    12.0    4.0
2         1     3         2     NaN    NaN
3         2     1         4    20.0    5.0
4         2     2         3     NaN    NaN
5         2     3         3    16.0    1.0

为什么这样更优

  • 向量化操作:避免逐行循环,Pandas向量化方法处理大数据集时效率提升显著
  • 避免警告:直接对DataFrame列赋值,不会触发SettingWithCopyWarning
  • 代码更简洁:逻辑清晰,可读性更强

内容的提问来源于stack exchange,提问作者A Doe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 02:45:29