跨文档词项矩阵行相似度计算及textstat_simil用法咨询
Hey there! Let's tackle your two questions one by one—they're both super common when working with document-term matrices (DTMs) in quanteda, so I’m glad you asked.
1. Can you use textstat_simil() to compare individual documents from Matrix A directly?
Absolutely! You don’t need to merge matrices first and then clean up irrelevant cells—that’s inefficient, especially with large datasets. textstat_simil() has built-in functionality to target specific documents, which lets you compare individual entries from Matrix A without extra steps. Here are two straightforward approaches:
Option 1: Target specific docs in a merged matrix (if you already have dtm3)
If dtm3 is your combined matrix of A + B, just use the target parameter to isolate documents from Matrix A. First, identify which rows belong to Matrix A (assuming your row names have a consistent label like doc_A_1, doc_A_2, etc.):
# Get row names for all docs in Matrix A a_docs <- rownames(dtm3)[grepl("doc_A_", rownames(dtm3))] # Loop through each doc in A to calculate similarity for (doc in a_docs) { # Calculate cosine similarity between the target A doc and all docs in dtm3 sim_results <- textstat_simil(dtm3, method = "cosine", target = doc) # If you only care about similarity to Matrix B docs, filter out A docs from results sim_results <- sim_results[, !colnames(sim_results) %in% a_docs] # Do something with the results (save, print, analyze) print(paste("Similarity scores for", doc, "vs Matrix B:")) print(sim_results) }
Option 2: Compare Matrix A docs directly to Matrix B (no merging needed)
Even better—skip merging entirely by ensuring both matrices share the same feature set, then compare individual rows of A to all rows of B:
# First, align the feature sets of dtm_A and dtm_B (critical for accurate similarity) dtm_A_aligned <- dfm_match(dtm_A, features = union(featnames(dtm_A), featnames(dtm_B))) dtm_B_aligned <- dfm_match(dtm_B, features = union(featnames(dtm_A), featnames(dtm_B))) # Loop through each document in Matrix A for (i in seq_len(nrow(dtm_A_aligned))) { # Calculate cosine similarity between the i-th doc in A and all docs in B sim_scores <- textstat_simil(dtm_A_aligned[i, ], dtm_B_aligned, method = "cosine") # Process results as needed print(paste("Similarity scores for A doc", i, "vs all B docs:")) print(sim_scores) }
This approach avoids loading a huge merged matrix into memory, which is way more efficient for large datasets.
2. How to fix file size bloat when using as.matrix(d1)?
The problem here is that quanteda’s dfm objects are sparse matrices (by default, stored as dgCMatrix from the Matrix package). Sparse matrices only store non-zero values, which saves massive amounts of space. When you run as.matrix(), you convert it to a dense matrix—this forces R to store every single zero value, hence the file size doubling (or worse, depending on how sparse your DTM is).
Here are 3 solutions to fix this:
Stick with sparse matrices whenever possible: Almost all quanteda functions (including
textstat_simil()) work directly withdfmobjects, so you rarely need to convert to a dense matrix. If you need to perform matrix operations outside quanteda, use the Matrix package’s sparse matrix methods instead. For example, to access a subset of the sparse matrix:# Access the first 5 rows of the sparse dfm (no conversion to dense matrix) sparse_subset <- dtm_A[1:5, ]Only convert subsets if you must: If you absolutely need a dense matrix, don’t convert the entire DTM—only the rows/columns you actually need. For example:
# Convert only the first 10 docs of dtm_A to a dense matrix dense_subset <- as.matrix(dtm_A[1:10, ])Save sparse matrices efficiently: If you’re saving your DTM to disk, use sparse-friendly formats instead of dense ones. Save the
dfmdirectly withsaveRDS():# Save the sparse dfm to disk (small file size) saveRDS(dtm_A, "dtm_A_sparse.rds") # Load it back later (still a sparse matrix) dtm_A_loaded <- readRDS("dtm_A_sparse.rds")You can also use the Matrix package’s
writeMM()to save in the Matrix Market format, which is designed for sparse data.
内容的提问来源于stack exchange,提问作者john

