高效查找包含numpy数组特定索引组合的行
Let's break this down into efficient, vectorized operations that avoid generating all possible index combinations—critical for large datasets. Here's how to tackle it step by step:
1. Extract Top 5 Indices for Each Column in np_weight
Instead of full sorting (which is slow for large arrays), use numpy.argpartition to get the top 5 indices per column in linear time. This is way more efficient than argsort when you only need the top N elements.
import numpy as np # Example data (replace with your actual arrays) np_weight = np.random.uniform(size=(7, 4)) # 7 features, 4 clusters/columns np_sentence = np.random.randint(0, 7, size=(5, 3)) # 5 sentences, 3 words each # Get top 5 indices for each column (since we have 7 rows, top 5 skips the bottom 2) top5_indices_per_col = np.argpartition(-np_weight, kth=5, axis=0)[:5, :] # Flatten to get all unique top indices (adjust if you need per-column checks) all_top5_indices = np.unique(top5_indices_per_col.flatten())
2. Flag Words in Sentences That Are in Top Indices
Create a boolean mask to mark which elements in np_sentence are part of our top indices. We use numpy.isin for vectorized checking—no slow Python loops here.
# Boolean array where True = word index is in our top set is_top_index = np.isin(np_sentence, all_top5_indices)
3. Filter Sentences with At Least Two Top Indices
Sum the boolean values along each row (since True counts as 1) and find rows where the sum is ≥2. This gives us the indices of sentences that meet your criteria.
# Count how many top indices are in each sentence top_count_per_sentence = is_top_index.sum(axis=1) # Get indices of sentences with ≥2 top indices matching_sentence_indices = np.where(top_count_per_sentence >= 2)[0] print("Matching sentence indices:", matching_sentence_indices)
If You Need Sentences with Two Indices from the Same Column's Top 5
If your requirement is stricter (sentences must contain two indices from the same column's top 5, not just any top indices), adjust the approach to check per column:
matching_indices = set() # Iterate over each column's top 5 indices for col_top5 in top5_indices_per_col.T: # Mask for this column's top indices col_mask = np.isin(np_sentence, col_top5) # Count matches per sentence col_count = col_mask.sum(axis=1) # Add qualifying sentences to our set matching_indices.update(np.where(col_count >= 2)[0]) # Convert to sorted numpy array matching_sentence_indices = np.array(sorted(matching_indices))
Why This Is Efficient
- Vectorized Operations: All core steps use numpy's C-implemented vectorized functions, which are orders of magnitude faster than Python loops for large datasets.
- No Combination Generation: Instead of creating every possible pair of top indices (which blows up in size as features grow), we count occurrences directly.
- Fast Top-N Extraction:
argpartitionruns in O(n) time per column, vs. O(n log n) for full sorting.
内容的提问来源于stack exchange,提问作者sariii

