如何将无重复Pandas数组转其他结构?优化10k列值去重排序性能
Hey there! Let's fix that slow, memory-heavy code of yours—handling 10k elements doesn't have to be a slog. Here's how to optimize your workflow and pick the right data structures for the job:
Optimized Pandas-Based Deduplication & Sorting
Your current loop uses Python's native list operations, which are slow for large datasets because checking if i not in arr runs in O(n) time per element (leading to an overall O(n²) time complexity). Pandas and NumPy have built-in vectorized operations that do this way faster:
Option 1: Get a sorted, deduplicated NumPy array
This is the most memory-efficient and fast option for basic use cases:
import pandas as pd df = pd.read_csv('path', sep=';') # Get unique values (returns a NumPy array) and sort it unique_sorted_arr = df.iloc[:, 0].unique() unique_sorted_arr.sort() # In-place sort (O(n log n) time)
df.iloc[:, 0] targets the first column (matching your original df[0]). The unique() method leverages NumPy's optimized C-backed logic to deduplicate in O(n) time, and sorting uses efficient low-level algorithms.
Option 2: Get a sorted, deduplicated Pandas Series
If you want to retain Pandas' built-in indexing and functionality (like easy filtering or alignment with other data), use this:
unique_sorted_series = df.iloc[:, 0].drop_duplicates().sort_values()
This returns a Series where values are deduplicated and sorted. You can quickly look up values using methods like unique_sorted_series.isin([target_value]) or direct boolean indexing.
Converting to Other Data Structures for Fast Lookups
Depending on your needs (like O(1) lookups or dynamic sorted operations), here are the best options:
1. NumPy Array (Recommended for Static Data)
The unique_sorted_arr from Option 1 is already a NumPy array. To enable fast index-based lookups, use the bisect module for binary search (O(log n) time per lookup):
import bisect target = "your_value_here" # Find the insertion point for the target in the sorted array idx = bisect.bisect_left(unique_sorted_arr, target) if idx < len(unique_sorted_arr) and unique_sorted_arr[idx] == target: print(f"Value found at index {idx}") else: print("Value not present")
2. Sorted List (For Dynamic Operations)
If you need to frequently add/remove values while keeping the list sorted and deduplicated, use the SortedList from the sortedcontainers library (install with pip install sortedcontainers):
from sortedcontainers import SortedList # Deduplicate first with a set, then create a sorted list unique_sorted_list = SortedList(set(df.iloc[:, 0].values)) # Lookups, inserts, and deletes all run in O(log n) time if target in unique_sorted_list: idx = unique_sorted_list.index(target)
3. Python Set (For O(1) Existence Checks)
If you only need to check if a value exists (and don't care about order), a set is unbeatable for speed:
unique_set = set(df.iloc[:, 0].values) if target in unique_set: print("Value exists")
Note: Sets are unordered, so if you need sorted results, convert to a sorted list afterward with sorted(unique_set).
Why This Is Better Than Your Original Code
- Speed: Pandas/NumPy operations are implemented in C, avoiding the overhead of Python loops. For 10k elements, this cuts runtime from seconds to milliseconds.
- Memory: NumPy arrays store values in contiguous memory blocks, which is far more efficient than Python lists (which store references to individual objects).
内容的提问来源于stack exchange,提问作者nexla

