如何用R从含冗余数据的3列CSV生成DTM矩阵?
Hey there! Let's break down why your DTM isn't shrinking and fix this properly.
First off, let's get to the root of the problem: the tm package's Corpus() and DTM tools are built for unstructured text corpora (think piles of articles or paragraphs). Your data is structured: it's a list of disease-gene-value associations, where redundancy comes from the same disease-gene pair popping up multiple times. Feeding each row directly into Corpus treats every single row as a separate "document"—so no wonder your DTM is still 600k rows big!
Here's the step-by-step solution to trim the redundancy and build the matrix you actually need:
1. Aggregate Redundant Disease-Gene Pairs First
We need to collapse duplicate disease-gene combinations by summarizing their associated values (use sum, mean, or another metric depending on your goals). For 600k rows, data.table is way faster than base R for this task:
# Load the fast data handling package library(data.table) # Read your CSV (replace with your file path) dt <- fread("your_data.csv", col.names = c("disease", "gene", "value")) # Group by disease + gene, sum the values (swap sum() with mean() if needed) aggregated_data <- dt[, .(total_value = sum(value)), by = .(disease, gene)]
This step will immediately cut down your rows to only unique disease-gene pairs—no more redundant entries cluttering your data.
2. Build Your Target Matrix (2 Options)
Now you can convert this cleaned data into a matrix that makes sense for your use case:
Option A: Wide-Format Disease-Gene Matrix
If you want a straightforward matrix where rows = diseases, columns = genes, and cells = aggregated values:
# Convert long-form aggregated data to wide format disease_gene_matrix <- dcast(aggregated_data, disease ~ gene, value.var = "total_value", fill = 0)
Missing disease-gene pairs will be filled with 0, and you'll end up with ~4000 rows (diseases) and ~16000 columns (genes).
Option B: Sparse DTM (For Memory-Efficient Large-Scale Data)
If you specifically need a Document-Term Matrix (treating diseases as "documents" and genes as "terms" with weighted values), use tm properly by first formatting each disease's genes/values as structured text:
library(tm) # For each disease, create a string like "gene1:value1 gene2:value2" disease_text <- aggregated_data[, .(text = paste(paste(gene, total_value, sep = ":"), collapse = " ")), by = disease] # Build a corpus from this structured text corpus <- VCorpus(DataframeSource(disease_text)) # Custom tokenizer to parse gene names and their weighted values weighted_tokenizer <- function(x) { tokens <- strsplit(x, " ")[[1]] lapply(tokens, function(t) { parts <- strsplit(t, ":")[[1]] list(term = parts[1], weight = as.numeric(parts[2])) }) } # Build the weighted DTM dtm <- DocumentTermMatrix(corpus, control = list(tokenize = weighted_tokenizer, weighting = function(x) weightTfIdf(x, normalize = FALSE)))
This DTM will have ~4000 rows (diseases) and ~16000 columns (genes)—exactly what you need, no more redundant entries.
Quick Pro Tip
Always clean and aggregate structured data before feeding it into text-processing tools like tm. Text tools aren't designed to handle duplicate structured pairs, so that step has to come first!
内容的提问来源于stack exchange,提问作者Jzk45

