You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将无重复Pandas数组转其他结构?优化10k列值去重排序性能

Hey there! Let's fix that slow, memory-heavy code of yours—handling 10k elements doesn't have to be a slog. Here's how to optimize your workflow and pick the right data structures for the job:

Optimized Pandas-Based Deduplication & Sorting

Your current loop uses Python's native list operations, which are slow for large datasets because checking if i not in arr runs in O(n) time per element (leading to an overall O(n²) time complexity). Pandas and NumPy have built-in vectorized operations that do this way faster:

Option 1: Get a sorted, deduplicated NumPy array

This is the most memory-efficient and fast option for basic use cases:

import pandas as pd

df = pd.read_csv('path', sep=';')
# Get unique values (returns a NumPy array) and sort it
unique_sorted_arr = df.iloc[:, 0].unique()
unique_sorted_arr.sort()  # In-place sort (O(n log n) time)

df.iloc[:, 0] targets the first column (matching your original df[0]). The unique() method leverages NumPy's optimized C-backed logic to deduplicate in O(n) time, and sorting uses efficient low-level algorithms.

Option 2: Get a sorted, deduplicated Pandas Series

If you want to retain Pandas' built-in indexing and functionality (like easy filtering or alignment with other data), use this:

unique_sorted_series = df.iloc[:, 0].drop_duplicates().sort_values()

This returns a Series where values are deduplicated and sorted. You can quickly look up values using methods like unique_sorted_series.isin([target_value]) or direct boolean indexing.

Converting to Other Data Structures for Fast Lookups

Depending on your needs (like O(1) lookups or dynamic sorted operations), here are the best options:

The unique_sorted_arr from Option 1 is already a NumPy array. To enable fast index-based lookups, use the bisect module for binary search (O(log n) time per lookup):

import bisect

target = "your_value_here"
# Find the insertion point for the target in the sorted array
idx = bisect.bisect_left(unique_sorted_arr, target)
if idx < len(unique_sorted_arr) and unique_sorted_arr[idx] == target:
    print(f"Value found at index {idx}")
else:
    print("Value not present")

2. Sorted List (For Dynamic Operations)

If you need to frequently add/remove values while keeping the list sorted and deduplicated, use the SortedList from the sortedcontainers library (install with pip install sortedcontainers):

from sortedcontainers import SortedList

# Deduplicate first with a set, then create a sorted list
unique_sorted_list = SortedList(set(df.iloc[:, 0].values))
# Lookups, inserts, and deletes all run in O(log n) time
if target in unique_sorted_list:
    idx = unique_sorted_list.index(target)

3. Python Set (For O(1) Existence Checks)

If you only need to check if a value exists (and don't care about order), a set is unbeatable for speed:

unique_set = set(df.iloc[:, 0].values)
if target in unique_set:
    print("Value exists")

Note: Sets are unordered, so if you need sorted results, convert to a sorted list afterward with sorted(unique_set).

Why This Is Better Than Your Original Code

  • Speed: Pandas/NumPy operations are implemented in C, avoiding the overhead of Python loops. For 10k elements, this cuts runtime from seconds to milliseconds.
  • Memory: NumPy arrays store values in contiguous memory blocks, which is far more efficient than Python lists (which store references to individual objects).

内容的提问来源于stack exchange,提问作者nexla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:45:20