Python如何判断DataFrame内的数值是否落在指定区间范围内
基于Pandas的实现方案
方法1:逐行匹配(适合小数据量场景)
逻辑简单直观,代码易读性高,适合数据量不大的场景使用:
import pandas as pd # 构造区间参考DataFrame df_interval = pd.DataFrame({ "Name": ["Blue", "Red", "Green", "Purple", "Yellow"], "Start": [10, 23, 89, 168, 21], "End": [28, 25, 107, 216, 40] }) # 构造待校验DataFrame df_check = pd.DataFrame({ "Name": ["W", "X", "Y", "Z"], "Value": [37, 176, 43, 96] }) contained = [] not_contained = [] for _, row in df_check.iterrows(): current_val = row["Value"] # 筛选符合区间条件的匹配项 match_res = df_interval[(df_interval["Start"] <= current_val) & (df_interval["End"] >= current_val)] if not match_res.empty: # 匹配成功加入contained,此处默认取第一个匹配区间,有多个匹配需求可自行调整逻辑 contained.append({ "check_name": row["Name"], "value": current_val, "matched_interval": match_res.iloc[0]["Name"] }) else: # 无匹配加入not_contained not_contained.append({ "check_name": row["Name"], "value": current_val }) print(contained) print(not_contained)
方法2:IntervalIndex优化(适合大数据量场景)
如果待校验数据或区间数量很大,用IntervalIndex可以大幅提升查询效率,避免逐行遍历的性能损耗:
# 构造左闭右闭的区间索引,绑定对应区间名称 interval_idx = pd.IntervalIndex.from_arrays(df_interval["Start"], df_interval["End"], closed="both") interval_map = pd.Series(df_interval["Name"].values, index=interval_idx) # 一次性完成所有值的匹配查询 df_check["matched_interval"] = df_check["Value"].map(interval_map) # 直接拆分得到两个结果列表 contained = df_check[df_check["matched_interval"].notna()].to_dict("records") not_contained = df_check[df_check["matched_interval"].isna()].drop(columns="matched_interval").to_dict("records")
内容的提问来源于stack exchange,提问作者Iacopo Passeri
相关产品推荐
相关产品推荐

