如何将含字符串键的文档-词二维字典转为支持矩阵乘法的矩阵?
Fixing scipy dok_matrix Conversion for Document Dictionary
It sounds like the core issue here is that scipy.sparse.dok_matrix relies on integer row/column indices, not the string keys (document IDs or words) you’re using in your dictionaries. Let’s walk through a complete, working solution to convert your document structure into a sparse matrix correctly.
Step-by-Step Implementation
First, we need to map your string labels to integer indices, then populate the matrix using those indices. Here’s a polished version of your function that fixes common pitfalls:
Full Working Code
import random from scipy.sparse import dok_matrix def getMatrix_N(trainDocs, N): # Sample N random document IDs from the outer dictionary docIds = random.sample(list(trainDocs.keys()), k=N) # Build a vocabulary mapping: word -> unique column index vocab = set() for doc_id in docIds: vocab.update(trainDocs[doc_id].keys()) vocab = sorted(vocab) # Optional, but ensures consistent index ordering word_to_idx = {word: idx for idx, word in enumerate(vocab)} num_columns = len(vocab) # Initialize the sparse matrix: N rows (sampled docs), num_columns columns (vocab) doc_matrix = dok_matrix((N, num_columns), dtype=float) # Populate the matrix with word values for row_idx, doc_id in enumerate(docIds): word_values = trainDocs[doc_id] for word, value in word_values.items(): col_idx = word_to_idx[word] doc_matrix[row_idx, col_idx] = value return doc_matrix
Key Fixes & Explanations
- String-to-index mapping: We create
word_to_idxto convert your word strings into integer column indices—this is the main missing piece in your original code. - Consistent matrix shape: The matrix is initialized with exactly N rows (one per sampled document) and columns equal to the number of unique words in your sampled docs.
- Avoiding KeyErrors: By building the vocabulary from the sampled documents directly, we ensure every word in those docs has a corresponding index.
Example Test Case
If your trainDocs looks like this:
trainDocs = { "recipe_1": {"flour": 0.8, "sugar": 0.5}, "recipe_2": {"eggs": 0.7, "flour": 0.3}, "recipe_3": {"sugar": 0.6, "eggs": 0.2} } # Generate matrix with 2 random documents sparse_mat = getMatrix_N(trainDocs, 2) print(sparse_mat.toarray()) # Convert to dense array to visualize
You’ll get a sparse matrix where each row represents a document, and each column represents a word with its corresponding float value.
内容的提问来源于stack exchange,提问作者Fancypants753
相关产品推荐
相关产品推荐

