使用pandas按多列阈值条件填充数据集observation空列的方法
数据集observation列填充最优实现方案
方案说明
最优实现采用pandas向量化条件判断,相比逐行遍历(apply/iterrows等方式)性能提升可达百倍以上,同时逻辑清晰易维护,既适合小样本数据集也兼容百万级以上的大规模数据集。
实现步骤
1. 依赖导入
首先导入需要的处理库:
import pandas as pd import numpy as np
2. 条件逻辑实现
核心逻辑通过两个布尔掩码匹配规则,再批量赋值:
- 阈值定义:
temp_thresh=40、humd_thresh=50、moist_thresh=150 - 匹配「所有指标均小于对应阈值」的行,赋值为
yes - 匹配「所有指标均大于对应阈值」的行,赋值为
no - 不满足上述两个条件的行,默认保留空值(可按需修改默认值)
完整代码示例:
# 构造样例数据集 df = pd.DataFrame({ "Temperature": [51,51,50,50,50], "Humidity": [29.5,29.5,29.5,29.5,29.5], "Moisture": [0,188,0,350,0], "observation": "" }) # 定义阈值 temp_thresh = 40 humd_thresh = 50 moist_thresh = 150 # 构造条件掩码 cond_yes = (df["Temperature"] < temp_thresh) & (df["Humidity"] < humd_thresh) & (df["Moisture"] < moist_thresh) cond_no = (df["Temperature"] > temp_thresh) & (df["Humidity"] > humd_thresh) & (df["Moisture"] > moist_thresh) # 批量赋值 df["observation"] = np.select([cond_yes, cond_no], ["yes", "no"], default="")
3. 样例输出说明
你提供的样例数据中所有行的Humidity均为29.5(小于50阈值)、Temperature均大于40阈值,没有行满足「全小于」或「全大于」的条件,因此最终样例的observation列均为空值。
内容的提问来源于stack exchange,提问作者AL HANA T
相关产品推荐
相关产品推荐

