You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP场景下如何高效计算区分左右窗口的正交性矩阵?

Hey there! Let's fix that efficiency bottleneck in your context matrix calculation. The main issue with your current code is that you're iterating over every word individually, which gets painfully slow with large corpora. Instead, we'll shift to processing all words in a single pass per sentence, then aggregate counts in bulk—this will cut down your runtime drastically.

First, Let's Break Down the Bottlenecks

  • Your contextDefiner runs once per vocabulary word, leading to an O(V*N) time complexity (V = vocabulary size, N = total words in corpus). For large datasets, this is a killer.
  • Filling a lil_matrix one cell at a time is inefficient; building a sparse matrix from pre-collected (row, column, value) triples is way faster.

Optimized Solution: One Pass Per Sentence

We'll rewrite the logic to extract all target-context pairs in a single pass for each sentence, then count pairs in bulk before building the sparse matrix.

Step 1: Optimized Vocabulary Creator (Minor Tweak)

First, simplify your vocabulary generator for cleaner code:

def vocabularyCreator(corpus):
    all_words = wordBreaker(corpus)
    vocabulary = sorted(set(all_words))
    # Use zip to map words to indices concisely
    return dict(zip(vocabulary, range(len(vocabulary))))

Step 2: Extract All Context Pairs in One Pass

This function processes a single tokenized sentence and returns all valid (target word index, context word index) pairs for either left or right windows:

def get_context_pairs(brokensentence, vocab, side):
    sent_len = len(brokensentence)
    pairs = []
    word_to_idx = vocab

    if side == 'right':
        # Target word at position i, context = i+1 to i+5
        for i in range(sent_len - 5):
            target_idx = word_to_idx.get(brokensentence[i])
            if target_idx is None:
                continue  # Skip words not in vocabulary
            # Add all 5 right-window words as pairs
            for j in range(1, 6):
                context_idx = word_to_idx.get(brokensentence[i+j])
                if context_idx is not None:
                    pairs.append((target_idx, context_idx))
    elif side == 'left':
        # Target word at position i, context = i-5 to i-1
        for i in range(5, sent_len):
            target_idx = word_to_idx.get(brokensentence[i])
            if target_idx is None:
                continue
            # Add all 5 left-window words as pairs
            for j in range(5):
                context_idx = word_to_idx.get(brokensentence[i-5 + j])
                if context_idx is not None:
                    pairs.append((target_idx, context_idx))
    return pairs

Step 3: Build Sparse Matrix in Bulk

Now we'll process the entire corpus, aggregate all pairs, count them, and build the CSR matrix in one go:

from collections import Counter
from scipy.sparse import csr_matrix
import numpy as np

def build_context_matrix(corpus_broken, vocab, side):
    # corpus_broken = list of tokenized sentences (each sentence is a list of words)
    all_pairs = []
    for sentence in corpus_broken:
        all_pairs.extend(get_context_pairs(sentence, vocab, side))
    
    # Count occurrences of each (target, context) pair
    pair_counts = Counter(all_pairs)
    
    # Extract rows, columns, and count values for CSR matrix
    rows, cols, data = [], [], []
    for (target_idx, context_idx), count in pair_counts.items():
        rows.append(target_idx)
        cols.append(context_idx)
        data.append(count)
    
    # Build the sparse matrix
    vocab_size = len(vocab)
    return csr_matrix((data, (rows, cols)), shape=(vocab_size, vocab_size))

How to Use This

# Assume your raw corpus is stored in `your_corpus`, and `wordBreaker` tokenizes it
vocab = vocabularyCreator(your_corpus)
# Tokenize your entire corpus first (once!)
corpus_broken = [wordBreaker(sentence) for sentence in your_corpus]

# Generate left and right context matrices
left_matrix = build_context_matrix(corpus_broken, vocab, 'left')
right_matrix = build_context_matrix(corpus_broken, vocab, 'right')

# Combine matrices if needed (e.g., add left and right counts)
combined_matrix = left_matrix + right_matrix

# This matrix works directly with sklearn's KMeans
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=5)
kmeans.fit(combined_matrix)

Extra Optimizations for Ultra-Large Corpora

If your corpus is massive (millions of sentences), you can avoid loading all pairs into memory at once:

  • Process sentences in batches (e.g., 1000 sentences at a time)
  • Update a global Counter or numpy-based count array incrementally
  • Use numpy.unique instead of Counter for pair counting (faster for huge datasets):
    pairs_np = np.array(all_pairs, dtype=int)
    unique_pairs, counts = np.unique(pairs_np, axis=0, return_counts=True)
    rows = unique_pairs[:, 0]
    cols = unique_pairs[:, 1]
    data = counts
    

内容的提问来源于stack exchange,提问作者Salwa Tabet Gonzalez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:41:12