You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

高效查找包含numpy数组特定索引组合的行

Efficiently Find Sentences Containing Multiple Top-Weighted Feature Indices

Let's break this down into efficient, vectorized operations that avoid generating all possible index combinations—critical for large datasets. Here's how to tackle it step by step:

1. Extract Top 5 Indices for Each Column in np_weight

Instead of full sorting (which is slow for large arrays), use numpy.argpartition to get the top 5 indices per column in linear time. This is way more efficient than argsort when you only need the top N elements.

import numpy as np

# Example data (replace with your actual arrays)
np_weight = np.random.uniform(size=(7, 4))  # 7 features, 4 clusters/columns
np_sentence = np.random.randint(0, 7, size=(5, 3))  # 5 sentences, 3 words each

# Get top 5 indices for each column (since we have 7 rows, top 5 skips the bottom 2)
top5_indices_per_col = np.argpartition(-np_weight, kth=5, axis=0)[:5, :]
# Flatten to get all unique top indices (adjust if you need per-column checks)
all_top5_indices = np.unique(top5_indices_per_col.flatten())

2. Flag Words in Sentences That Are in Top Indices

Create a boolean mask to mark which elements in np_sentence are part of our top indices. We use numpy.isin for vectorized checking—no slow Python loops here.

# Boolean array where True = word index is in our top set
is_top_index = np.isin(np_sentence, all_top5_indices)

3. Filter Sentences with At Least Two Top Indices

Sum the boolean values along each row (since True counts as 1) and find rows where the sum is ≥2. This gives us the indices of sentences that meet your criteria.

# Count how many top indices are in each sentence
top_count_per_sentence = is_top_index.sum(axis=1)
# Get indices of sentences with ≥2 top indices
matching_sentence_indices = np.where(top_count_per_sentence >= 2)[0]

print("Matching sentence indices:", matching_sentence_indices)

If You Need Sentences with Two Indices from the Same Column's Top 5

If your requirement is stricter (sentences must contain two indices from the same column's top 5, not just any top indices), adjust the approach to check per column:

matching_indices = set()

# Iterate over each column's top 5 indices
for col_top5 in top5_indices_per_col.T:
    # Mask for this column's top indices
    col_mask = np.isin(np_sentence, col_top5)
    # Count matches per sentence
    col_count = col_mask.sum(axis=1)
    # Add qualifying sentences to our set
    matching_indices.update(np.where(col_count >= 2)[0])

# Convert to sorted numpy array
matching_sentence_indices = np.array(sorted(matching_indices))

Why This Is Efficient

  • Vectorized Operations: All core steps use numpy's C-implemented vectorized functions, which are orders of magnitude faster than Python loops for large datasets.
  • No Combination Generation: Instead of creating every possible pair of top indices (which blows up in size as features grow), we count occurrences directly.
  • Fast Top-N Extraction: argpartition runs in O(n) time per column, vs. O(n log n) for full sorting.

内容的提问来源于stack exchange,提问作者sariii

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:11:53