优化Pandas列中prediction1与符合条件的prediction2索引对的查找方法
优化Pandas列中prediction1与符合条件的prediction2索引对的查找方法
嗨,我明白你现在的问题——用嵌套的iterrows循环来找这些索引对确实效率很低,尤其是当你的DataFrame数据量变大的时候,这种方法会慢得让人头疼。下面我给你分享一个更高效的方法,利用Pandas和Numpy的向量化操作来解决这个问题:
核心思路
我们可以先把所有符合条件的索引单独提取出来,再通过二分查找快速为每个prediction1匹配到它之后第一个符合要求的prediction2,完全避开嵌套循环的低效操作:
- 提取所有
prediction为prediction1的索引列表 - 提取所有
prediction为prediction2且score≥0.85的索引列表(这个列表本身是按索引递增排序的) - 对每个
prediction1的索引,用二分查找找到它之后第一个符合条件的prediction2索引,形成配对
具体代码实现
import pandas as pd import numpy as np # 先还原你的DataFrame data = ['others', 'others', 'prediction1', 'others', 'others', 'others', 'others', 'others', 'others', 'others', 'prediction2', 'others', 'prediction2', 'others', 'others', 'prediction1', 'others', 'others', 'others', 'others', 'others', 'others', 'others', 'prediction2', 'others', 'others', 'others', 'others', 'prediction1', 'others'] score = [0.75, 0.75, 0.9, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.88, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75, 0.9, 0.75, 0.75, 0.75, 0.75, 0.75, 0.75] df = pd.DataFrame({'prediction': data, 'score': score}) # 步骤1:获取所有prediction1的索引 pred1_indices = df[df['prediction'] == 'prediction1'].index.tolist() # 步骤2:获取所有符合条件的prediction2的索引(score≥0.85) valid_pred2_indices = df[(df['prediction'] == 'prediction2') & (df['score'] >= 0.85)].index.tolist() # 步骤3:用二分查找匹配每个pred1对应的第一个有效pred2 pred_pairs = [] for pred1_idx in pred1_indices: # 找到第一个大于pred1_idx的valid_pred2索引的位置 pos = np.searchsorted(valid_pred2_indices, pred1_idx, side='right') # 如果这个位置在有效列表范围内,说明找到匹配项 if pos < len(valid_pred2_indices): pred_pairs.append((pred1_idx, valid_pred2_indices[pos])) print(pred_pairs) # 输出:[(2, 12), (15, 23)]
为什么这个方法更快?
- 原来的嵌套
iterrows是**O(n²)**的时间复杂度,数据量越大,耗时增长得越快; - 新方法用了向量化筛选和二分查找,时间复杂度降到O(n log m)(n是prediction1的数量,m是有效prediction2的数量),在大数据量下效率提升非常明显。
额外优化提示
如果你的DataFrame是非常大的数据集,还可以进一步提速:把valid_pred2_indices转换成Numpy数组,np.searchsorted处理数组的速度会比列表更快,代码只需要修改这一行:
valid_pred2_indices = df[(df['prediction'] == 'prediction2') & (df['score'] >= 0.85)].index.to_numpy()
备注:内容来源于stack exchange,提问作者Akshit jain
相关产品推荐
相关产品推荐

