You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解读sklearn稀疏矩阵输出?二元词共现矩阵构建疑问

Fixing Your Bigram Co-occurrence Matrix Issue

Hey there! Let’s work through this bigram co-occurrence matrix problem together. It sounds like you’re aiming for a matrix where each row is a preceding word, each column is a following word, and the value at (i,j) counts how many times word i is directly followed by word j. If your current output from Xc (and even Xc.todense()) is unreadable or not matching what you expect, let’s break down what might be going wrong and how to build the matrix correctly.

Common Pitfalls to Check First

First, let’s rule out easy mistakes:

  • Did you mix up row/column meaning? Some default co-occurrence implementations count column words preceding row words, which would flip your intended matrix.
  • Are you using a sliding window instead of strict bigrams? If your code uses a window larger than 1, it’ll count non-adjacent words too, which isn’t a true bigram follow-count matrix.
  • Did you skip proper preprocessing? Unnormalized text (mixed case, punctuation, extra spaces) can create duplicate entries in your vocabulary, making the matrix look messy.

Step 1: Build the Matrix Manually (Great for Understanding)

If you want to see exactly how the counts are generated, start with a manual implementation for small corpora. This makes it easy to verify each step:

import numpy as np
from collections import defaultdict

# Example corpus (replace with your own)
corpus = [
    "I love natural language processing",
    "I love coding in python",
    "natural language processing is fun"
]

# Preprocess: tokenize, lowercase, remove punctuation (adjust as needed)
tokenized_corpus = [sent.lower().split() for sent in corpus]

# Create a sorted vocabulary and word-to-index mapping
all_tokens = [tok for sent in tokenized_corpus for tok in sent]
vocab = sorted(set(all_tokens))
word_to_idx = {word: idx for idx, word in enumerate(vocab)}
vocab_size = len(vocab)

# Initialize empty bigram matrix (rows = preceding words, columns = following words)
bigram_matrix = np.zeros((vocab_size, vocab_size), dtype=int)

# Populate the matrix by counting adjacent word pairs
for sent in tokenized_corpus:
    for i in range(len(sent) - 1):
        prev_word = sent[i]
        next_word = sent[i+1]
        prev_idx = word_to_idx[prev_word]
        next_idx = word_to_idx[next_word]
        bigram_matrix[prev_idx][next_idx] += 1

# Print results to verify
print("Vocabulary:", vocab)
print("\nBigram Follow-Count Matrix:")
print(bigram_matrix)

When you run this, you’ll get a matrix where, for example, the row for "love" will have counts for the words that directly follow it ("natural" and "coding").

Step 2: Efficient Implementation for Large Corpora

For bigger datasets, use sparse matrices to save memory. Here’s how to leverage scikit-learn and scipy to build the matrix efficiently:

from sklearn.feature_extraction.text import CountVectorizer
import scipy.sparse as sp

# Use your corpus here
corpus = [
    "I love natural language processing",
    "I love coding in python",
    "natural language processing is fun"
]

# Step 1: Get single-word vocabulary and mapping
vectorizer = CountVectorizer(lowercase=True)
X_single = vectorizer.fit_transform(corpus)
vocab = vectorizer.get_feature_names_out()
word_to_idx = {word: idx for idx, word in enumerate(vocab)}

# Step 2: Count all bigrams
vectorizer_bigram = CountVectorizer(ngram_range=(2, 2), lowercase=True)
X_bigram = vectorizer_bigram.fit_transform(corpus)
bigram_vocab = vectorizer_bigram.get_feature_names_out()
bigram_counts = X_bigram.sum(axis=0).tolist()[0]

# Step 3: Build sparse bigram follow-count matrix
bigram_matrix = sp.lil_matrix((len(vocab), len(vocab)), dtype=int)
for bigram, count in zip(bigram_vocab, bigram_counts):
    prev_word, next_word = bigram.split()
    prev_idx = word_to_idx[prev_word]
    next_idx = word_to_idx[next_word]
    bigram_matrix[prev_idx, next_idx] = count

# Convert to dense only if your vocabulary is small!
print("\nDense Bigram Matrix:")
print(bigram_matrix.todense())

Why Your Original Xc Might Be Unreadable

If your original Xc was a sparse matrix from a library method, it might:

  • Be stored in a compressed format (like COO or CSR) that doesn’t print nicely without converting to dense.
  • Have rows/columns in an unexpected order (always check your vocabulary mapping to confirm which index corresponds to which word).
  • Include non-adjacent co-occurrences if you used a window-based co-occurrence method instead of strict bigrams.

内容的提问来源于stack exchange,提问作者quanty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:42:36