如何解读sklearn稀疏矩阵输出?二元词共现矩阵构建疑问
Hey there! Let’s work through this bigram co-occurrence matrix problem together. It sounds like you’re aiming for a matrix where each row is a preceding word, each column is a following word, and the value at (i,j) counts how many times word i is directly followed by word j. If your current output from Xc (and even Xc.todense()) is unreadable or not matching what you expect, let’s break down what might be going wrong and how to build the matrix correctly.
Common Pitfalls to Check First
First, let’s rule out easy mistakes:
- Did you mix up row/column meaning? Some default co-occurrence implementations count column words preceding row words, which would flip your intended matrix.
- Are you using a sliding window instead of strict bigrams? If your code uses a window larger than 1, it’ll count non-adjacent words too, which isn’t a true bigram follow-count matrix.
- Did you skip proper preprocessing? Unnormalized text (mixed case, punctuation, extra spaces) can create duplicate entries in your vocabulary, making the matrix look messy.
Step 1: Build the Matrix Manually (Great for Understanding)
If you want to see exactly how the counts are generated, start with a manual implementation for small corpora. This makes it easy to verify each step:
import numpy as np from collections import defaultdict # Example corpus (replace with your own) corpus = [ "I love natural language processing", "I love coding in python", "natural language processing is fun" ] # Preprocess: tokenize, lowercase, remove punctuation (adjust as needed) tokenized_corpus = [sent.lower().split() for sent in corpus] # Create a sorted vocabulary and word-to-index mapping all_tokens = [tok for sent in tokenized_corpus for tok in sent] vocab = sorted(set(all_tokens)) word_to_idx = {word: idx for idx, word in enumerate(vocab)} vocab_size = len(vocab) # Initialize empty bigram matrix (rows = preceding words, columns = following words) bigram_matrix = np.zeros((vocab_size, vocab_size), dtype=int) # Populate the matrix by counting adjacent word pairs for sent in tokenized_corpus: for i in range(len(sent) - 1): prev_word = sent[i] next_word = sent[i+1] prev_idx = word_to_idx[prev_word] next_idx = word_to_idx[next_word] bigram_matrix[prev_idx][next_idx] += 1 # Print results to verify print("Vocabulary:", vocab) print("\nBigram Follow-Count Matrix:") print(bigram_matrix)
When you run this, you’ll get a matrix where, for example, the row for "love" will have counts for the words that directly follow it ("natural" and "coding").
Step 2: Efficient Implementation for Large Corpora
For bigger datasets, use sparse matrices to save memory. Here’s how to leverage scikit-learn and scipy to build the matrix efficiently:
from sklearn.feature_extraction.text import CountVectorizer import scipy.sparse as sp # Use your corpus here corpus = [ "I love natural language processing", "I love coding in python", "natural language processing is fun" ] # Step 1: Get single-word vocabulary and mapping vectorizer = CountVectorizer(lowercase=True) X_single = vectorizer.fit_transform(corpus) vocab = vectorizer.get_feature_names_out() word_to_idx = {word: idx for idx, word in enumerate(vocab)} # Step 2: Count all bigrams vectorizer_bigram = CountVectorizer(ngram_range=(2, 2), lowercase=True) X_bigram = vectorizer_bigram.fit_transform(corpus) bigram_vocab = vectorizer_bigram.get_feature_names_out() bigram_counts = X_bigram.sum(axis=0).tolist()[0] # Step 3: Build sparse bigram follow-count matrix bigram_matrix = sp.lil_matrix((len(vocab), len(vocab)), dtype=int) for bigram, count in zip(bigram_vocab, bigram_counts): prev_word, next_word = bigram.split() prev_idx = word_to_idx[prev_word] next_idx = word_to_idx[next_word] bigram_matrix[prev_idx, next_idx] = count # Convert to dense only if your vocabulary is small! print("\nDense Bigram Matrix:") print(bigram_matrix.todense())
Why Your Original Xc Might Be Unreadable
If your original Xc was a sparse matrix from a library method, it might:
- Be stored in a compressed format (like COO or CSR) that doesn’t print nicely without converting to dense.
- Have rows/columns in an unexpected order (always check your vocabulary mapping to confirm which index corresponds to which word).
- Include non-adjacent co-occurrences if you used a window-based co-occurrence method instead of strict bigrams.
内容的提问来源于stack exchange,提问作者quanty

