You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

优化Pandas列中prediction1与符合条件的prediction2索引对的查找方法

优化Pandas列中prediction1与符合条件的prediction2索引对的查找方法

嗨,我明白你现在的问题——用嵌套的iterrows循环来找这些索引对确实效率很低,尤其是当你的DataFrame数据量变大的时候,这种方法会慢得让人头疼。下面我给你分享一个更高效的方法,利用Pandas和Numpy的向量化操作来解决这个问题:

核心思路

我们可以先把所有符合条件的索引单独提取出来,再通过二分查找快速为每个prediction1匹配到它之后第一个符合要求的prediction2,完全避开嵌套循环的低效操作:

  1. 提取所有prediction为prediction1的索引列表
  2. 提取所有prediction为prediction2且score≥0.85的索引列表(这个列表本身是按索引递增排序的)
  3. 对每个prediction1的索引,用二分查找找到它之后第一个符合条件的prediction2索引,形成配对

具体代码实现

import pandas as pd
import numpy as np

# 先还原你的DataFrame
data = ['others', 'others', 'prediction1', 'others', 'others', 'others', 'others', 'others', 'others', 'others', 'prediction2', 'others', 'prediction2', 'others', 'others', 'prediction1', 'others', 'others', 'others', 'others', 'others', 'others', 'others', 'prediction2', 'others', 'others', 'others', 'others', 'prediction1', 'others']
score = [0.75, 0.75, 0.9, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.88, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.9, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75]
df = pd.DataFrame({'prediction': data, 'score': score})

# 步骤1:获取所有prediction1的索引
pred1_indices = df[df['prediction'] == 'prediction1'].index.tolist()

# 步骤2:获取所有符合条件的prediction2的索引(score≥0.85)
valid_pred2_indices = df[(df['prediction'] == 'prediction2') & (df['score'] >= 0.85)].index.tolist()

# 步骤3:用二分查找匹配每个pred1对应的第一个有效pred2
pred_pairs = []
for pred1_idx in pred1_indices:
    # 找到第一个大于pred1_idx的valid_pred2索引的位置
    pos = np.searchsorted(valid_pred2_indices, pred1_idx, side='right')
    # 如果这个位置在有效列表范围内,说明找到匹配项
    if pos < len(valid_pred2_indices):
        pred_pairs.append((pred1_idx, valid_pred2_indices[pos]))

print(pred_pairs)
# 输出:[(2, 12), (15, 23)]

为什么这个方法更快?

  • 原来的嵌套iterrows是**O(n²)**的时间复杂度,数据量越大,耗时增长得越快;
  • 新方法用了向量化筛选和二分查找,时间复杂度降到O(n log m)(n是prediction1的数量,m是有效prediction2的数量),在大数据量下效率提升非常明显。

额外优化提示

如果你的DataFrame是非常大的数据集,还可以进一步提速:把valid_pred2_indices转换成Numpy数组,np.searchsorted处理数组的速度会比列表更快,代码只需要修改这一行:

valid_pred2_indices = df[(df['prediction'] == 'prediction2') & (df['score'] >= 0.85)].index.to_numpy()

备注:内容来源于stack exchange,提问作者Akshit jain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 14:23:03