You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于指定列值对NumPy数组中的列表进行重排分组?

Optimized Ways to Group NumPy Array by Second Column

Hey there! Your current approach gets the job done, but it's pretty inefficient because it iterates through the entire dataset 8000 times—that's a ton of redundant work, especially as your dataset grows. Let's fix that with two far more efficient methods:

Method 1: Pure NumPy (Sort + Split)

This approach leverages NumPy's optimized sorting and array splitting functions, which run in O(n log n) time (way better than your original O(n*k) complexity, where k is your loop count).

import numpy as np

# Load your data as before
my_data = np.genfromtxt(filename, delimiter=",", dtype=None)

# Step 1: Extract the second column as our grouping key
group_keys = my_data[:, 1]

# Step 2: Get indices that would sort the group keys
sorted_indices = np.argsort(group_keys)

# Step 3: Sort the original data using these indices
sorted_data = my_data[sorted_indices]

# Step 4: Find positions where the group key changes (split points)
sorted_keys = group_keys[sorted_indices]
split_points = np.where(sorted_keys[1:] != sorted_keys[:-1])[0] + 1

# Step 5: Split the sorted array into groups
grouped_arrays = np.split(sorted_data, split_points)

# Convert to the list format you need
result = [group.tolist() for group in grouped_arrays]

Method 2: Pandas (Simpler, Just as Efficient)

If you're open to using Pandas, its built-in groupby method is highly optimized and requires way less code:

import pandas as pd
import numpy as np

my_data = np.genfromtxt(filename, delimiter=",", dtype=None)

# Convert NumPy array to DataFrame
df = pd.DataFrame(my_data)

# Group by the second column (index 1) and convert each group to a list
result = [group.values.tolist() for _, group in df.groupby(1)]

Why Your Original Approach Is Slow

Your loop runs 8000 times, and every iteration scans the entire dataset to filter rows matching the current x value. For large datasets, this means you're doing the same work thousands of times over. Both methods above avoid this redundancy:

  • The NumPy method sorts once, then splits the sorted array in a single pass.
  • Pandas' groupby uses efficient hashing or sorting under the hood to group rows in one go.

Either of these methods will be drastically faster, especially as your dataset size increases.

内容的提问来源于stack exchange,提问作者Arsenal Fanatic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:17:57