You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将含字符串键的文档-词二维字典转为支持矩阵乘法的矩阵?

Fixing scipy dok_matrix Conversion for Document Dictionary

It sounds like the core issue here is that scipy.sparse.dok_matrix relies on integer row/column indices, not the string keys (document IDs or words) you’re using in your dictionaries. Let’s walk through a complete, working solution to convert your document structure into a sparse matrix correctly.

Step-by-Step Implementation

First, we need to map your string labels to integer indices, then populate the matrix using those indices. Here’s a polished version of your function that fixes common pitfalls:

Full Working Code

import random
from scipy.sparse import dok_matrix

def getMatrix_N(trainDocs, N):
    # Sample N random document IDs from the outer dictionary
    docIds = random.sample(list(trainDocs.keys()), k=N)
    
    # Build a vocabulary mapping: word -> unique column index
    vocab = set()
    for doc_id in docIds:
        vocab.update(trainDocs[doc_id].keys())
    vocab = sorted(vocab)  # Optional, but ensures consistent index ordering
    word_to_idx = {word: idx for idx, word in enumerate(vocab)}
    num_columns = len(vocab)
    
    # Initialize the sparse matrix: N rows (sampled docs), num_columns columns (vocab)
    doc_matrix = dok_matrix((N, num_columns), dtype=float)
    
    # Populate the matrix with word values
    for row_idx, doc_id in enumerate(docIds):
        word_values = trainDocs[doc_id]
        for word, value in word_values.items():
            col_idx = word_to_idx[word]
            doc_matrix[row_idx, col_idx] = value
    
    return doc_matrix

Key Fixes & Explanations

  • String-to-index mapping: We create word_to_idx to convert your word strings into integer column indices—this is the main missing piece in your original code.
  • Consistent matrix shape: The matrix is initialized with exactly N rows (one per sampled document) and columns equal to the number of unique words in your sampled docs.
  • Avoiding KeyErrors: By building the vocabulary from the sampled documents directly, we ensure every word in those docs has a corresponding index.

Example Test Case

If your trainDocs looks like this:

trainDocs = {
    "recipe_1": {"flour": 0.8, "sugar": 0.5},
    "recipe_2": {"eggs": 0.7, "flour": 0.3},
    "recipe_3": {"sugar": 0.6, "eggs": 0.2}
}

# Generate matrix with 2 random documents
sparse_mat = getMatrix_N(trainDocs, 2)
print(sparse_mat.toarray())  # Convert to dense array to visualize

You’ll get a sparse matrix where each row represents a document, and each column represents a word with its corresponding float value.

内容的提问来源于stack exchange,提问作者Fancypants753

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:12:54