You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas:含重复值的已排序Code列,二分查找vs直接匹配选哪个?

Hey there! Let's tackle your question step by step.

First, let's talk about your data size

With only ~500 rows, the performance difference between using bisect and a direct Pandas filter is completely negligible. You won't notice any speed gap here—so the biggest factor should be code simplicity and readability.

The simplest (and best) approach for your case: Pandas boolean indexing

Forget overcomplicating things with bisect here. Pandas has a built-in, optimized way to filter rows that match your target code, and it handles duplicates perfectly out of the box. Just use boolean indexing:

target_code = "PA0000359"
matching_rows = data[data['Code'] == target_code]

This code is straightforward, easy to maintain, and leverages Pandas' vectorized operations under the hood—even with 500 rows, it'll run in a blink.

What about bisect? When would it make sense?

If you were working with millions of rows, bisect's binary search could offer a speed boost over a linear scan (though Pandas' optimized filters are still pretty fast). But for 500 rows, it's overkill.

That said, if you still want to explore how bisect handles duplicates: since your Code column is sorted, you can use bisect_left to find the first occurrence of your target code, and bisect_right to find the position right after the last occurrence. Then slice your DataFrame between those indices to get all matching rows:

import bisect

# Extract the sorted Code column as a list
codes_list = data['Code'].tolist()
target_code = "PA0000359"

# Find left and right boundaries of the target code
left_pos = bisect.bisect_left(codes_list, target_code)
right_pos = bisect.bisect_right(codes_list, target_code)

# Slice the DataFrame to get all matching rows
matching_rows = data.iloc[left_pos:right_pos]

This works correctly with duplicates—bisect_left gives the first index where the target can be inserted to keep the list sorted (i.e., the first occurrence of the target), and bisect_right gives the index after the last occurrence.

Final verdict

Stick with the Pandas boolean indexing approach. It's simpler, more readable, and for your small dataset, there's no reason to use bisect. Reserve bisect for cases where you're dealing with extremely large sorted datasets and need the binary search performance edge.

内容的提问来源于stack exchange,提问作者Usman Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:09:09