NLP场景下如何高效计算区分左右窗口的正交性矩阵?
Hey there! Let's fix that efficiency bottleneck in your context matrix calculation. The main issue with your current code is that you're iterating over every word individually, which gets painfully slow with large corpora. Instead, we'll shift to processing all words in a single pass per sentence, then aggregate counts in bulk—this will cut down your runtime drastically.
First, Let's Break Down the Bottlenecks
- Your
contextDefinerruns once per vocabulary word, leading to an O(V*N) time complexity (V = vocabulary size, N = total words in corpus). For large datasets, this is a killer. - Filling a
lil_matrixone cell at a time is inefficient; building a sparse matrix from pre-collected (row, column, value) triples is way faster.
Optimized Solution: One Pass Per Sentence
We'll rewrite the logic to extract all target-context pairs in a single pass for each sentence, then count pairs in bulk before building the sparse matrix.
Step 1: Optimized Vocabulary Creator (Minor Tweak)
First, simplify your vocabulary generator for cleaner code:
def vocabularyCreator(corpus): all_words = wordBreaker(corpus) vocabulary = sorted(set(all_words)) # Use zip to map words to indices concisely return dict(zip(vocabulary, range(len(vocabulary))))
Step 2: Extract All Context Pairs in One Pass
This function processes a single tokenized sentence and returns all valid (target word index, context word index) pairs for either left or right windows:
def get_context_pairs(brokensentence, vocab, side): sent_len = len(brokensentence) pairs = [] word_to_idx = vocab if side == 'right': # Target word at position i, context = i+1 to i+5 for i in range(sent_len - 5): target_idx = word_to_idx.get(brokensentence[i]) if target_idx is None: continue # Skip words not in vocabulary # Add all 5 right-window words as pairs for j in range(1, 6): context_idx = word_to_idx.get(brokensentence[i+j]) if context_idx is not None: pairs.append((target_idx, context_idx)) elif side == 'left': # Target word at position i, context = i-5 to i-1 for i in range(5, sent_len): target_idx = word_to_idx.get(brokensentence[i]) if target_idx is None: continue # Add all 5 left-window words as pairs for j in range(5): context_idx = word_to_idx.get(brokensentence[i-5 + j]) if context_idx is not None: pairs.append((target_idx, context_idx)) return pairs
Step 3: Build Sparse Matrix in Bulk
Now we'll process the entire corpus, aggregate all pairs, count them, and build the CSR matrix in one go:
from collections import Counter from scipy.sparse import csr_matrix import numpy as np def build_context_matrix(corpus_broken, vocab, side): # corpus_broken = list of tokenized sentences (each sentence is a list of words) all_pairs = [] for sentence in corpus_broken: all_pairs.extend(get_context_pairs(sentence, vocab, side)) # Count occurrences of each (target, context) pair pair_counts = Counter(all_pairs) # Extract rows, columns, and count values for CSR matrix rows, cols, data = [], [], [] for (target_idx, context_idx), count in pair_counts.items(): rows.append(target_idx) cols.append(context_idx) data.append(count) # Build the sparse matrix vocab_size = len(vocab) return csr_matrix((data, (rows, cols)), shape=(vocab_size, vocab_size))
How to Use This
# Assume your raw corpus is stored in `your_corpus`, and `wordBreaker` tokenizes it vocab = vocabularyCreator(your_corpus) # Tokenize your entire corpus first (once!) corpus_broken = [wordBreaker(sentence) for sentence in your_corpus] # Generate left and right context matrices left_matrix = build_context_matrix(corpus_broken, vocab, 'left') right_matrix = build_context_matrix(corpus_broken, vocab, 'right') # Combine matrices if needed (e.g., add left and right counts) combined_matrix = left_matrix + right_matrix # This matrix works directly with sklearn's KMeans from sklearn.cluster import KMeans kmeans = KMeans(n_clusters=5) kmeans.fit(combined_matrix)
Extra Optimizations for Ultra-Large Corpora
If your corpus is massive (millions of sentences), you can avoid loading all pairs into memory at once:
- Process sentences in batches (e.g., 1000 sentences at a time)
- Update a global
Counteror numpy-based count array incrementally - Use
numpy.uniqueinstead ofCounterfor pair counting (faster for huge datasets):pairs_np = np.array(all_pairs, dtype=int) unique_pairs, counts = np.unique(pairs_np, axis=0, return_counts=True) rows = unique_pairs[:, 0] cols = unique_pairs[:, 1] data = counts
内容的提问来源于stack exchange,提问作者Salwa Tabet Gonzalez

