You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中高效批量获取first_valid_index的方法问询

高效实现有序event_time列的多阈值first_valid_index查询

因为你的event_time列是递增有序的,完全可以利用二分查找替代每次从头扫描的低效操作,以下是具体实现思路:

核心思路

原方法每次调用df.where(df.event_time >= 阈值).first_valid_index()都会遍历整个列做比较,时间复杂度为O(n);而针对有序序列的二分查找时间复杂度是O(logn),查询次数越多,效率提升越明显。

具体实现

单阈值查询

直接使用pandas.Series.searchsorted方法定位第一个满足event_time >= 阈值的位置,再映射回DataFrame的索引:

# 提前取出event_time列,避免重复调用
event_times = df['event_time']
threshold = 83000000

# 用searchsorted找第一个>=阈值的位置,side='left'确保返回左边界
pos = event_times.searchsorted(threshold, side='left')

# 处理边界情况:如果所有值都小于阈值,pos会等于df的行数
first_idx = df.index[pos] if pos < len(df) else None

这个结果和原方法df.where(df.event_time >= threshold).first_valid_index()完全一致,但速度快得多。

多阈值批量查询

如果需要查询多个阈值,可以一次性传入阈值列表,批量获取位置后再生成对应索引,效率进一步提升:

thresholds = [83000000, 84000000, 85000000]
positions = event_times.searchsorted(thresholds, side='left')

# 批量生成结果
first_indices = [
    df.index[p] if p < len(df) else None
    for p in positions
]

注意事项

  • 该方法仅适用于event_time列严格递增或非递减的场景,这正好符合你的前提条件;
  • 如果event_time存在重复值,side='left'依然会返回第一个出现的符合条件的位置,和原方法逻辑完全对齐。

内容的提问来源于stack exchange,提问作者En-Jui Kuo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 15:25:34