You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R从含冗余数据的3列CSV生成DTM矩阵?

Hey there! Let's break down why your DTM isn't shrinking and fix this properly.

Fixing Redundant Data & Building the Right Matrix

First off, let's get to the root of the problem: the tm package's Corpus() and DTM tools are built for unstructured text corpora (think piles of articles or paragraphs). Your data is structured: it's a list of disease-gene-value associations, where redundancy comes from the same disease-gene pair popping up multiple times. Feeding each row directly into Corpus treats every single row as a separate "document"—so no wonder your DTM is still 600k rows big!

Here's the step-by-step solution to trim the redundancy and build the matrix you actually need:

1. Aggregate Redundant Disease-Gene Pairs First

We need to collapse duplicate disease-gene combinations by summarizing their associated values (use sum, mean, or another metric depending on your goals). For 600k rows, data.table is way faster than base R for this task:

# Load the fast data handling package
library(data.table)

# Read your CSV (replace with your file path)
dt <- fread("your_data.csv", col.names = c("disease", "gene", "value"))

# Group by disease + gene, sum the values (swap sum() with mean() if needed)
aggregated_data <- dt[, .(total_value = sum(value)), by = .(disease, gene)]

This step will immediately cut down your rows to only unique disease-gene pairs—no more redundant entries cluttering your data.

2. Build Your Target Matrix (2 Options)

Now you can convert this cleaned data into a matrix that makes sense for your use case:

Option A: Wide-Format Disease-Gene Matrix

If you want a straightforward matrix where rows = diseases, columns = genes, and cells = aggregated values:

# Convert long-form aggregated data to wide format
disease_gene_matrix <- dcast(aggregated_data, disease ~ gene, value.var = "total_value", fill = 0)

Missing disease-gene pairs will be filled with 0, and you'll end up with ~4000 rows (diseases) and ~16000 columns (genes).

Option B: Sparse DTM (For Memory-Efficient Large-Scale Data)

If you specifically need a Document-Term Matrix (treating diseases as "documents" and genes as "terms" with weighted values), use tm properly by first formatting each disease's genes/values as structured text:

library(tm)

# For each disease, create a string like "gene1:value1 gene2:value2"
disease_text <- aggregated_data[, .(text = paste(paste(gene, total_value, sep = ":"), collapse = " ")), by = disease]

# Build a corpus from this structured text
corpus <- VCorpus(DataframeSource(disease_text))

# Custom tokenizer to parse gene names and their weighted values
weighted_tokenizer <- function(x) {
  tokens <- strsplit(x, " ")[[1]]
  lapply(tokens, function(t) {
    parts <- strsplit(t, ":")[[1]]
    list(term = parts[1], weight = as.numeric(parts[2]))
  })
}

# Build the weighted DTM
dtm <- DocumentTermMatrix(corpus, control = list(tokenize = weighted_tokenizer, weighting = function(x) weightTfIdf(x, normalize = FALSE)))

This DTM will have ~4000 rows (diseases) and ~16000 columns (genes)—exactly what you need, no more redundant entries.

Quick Pro Tip

Always clean and aggregate structured data before feeding it into text-processing tools like tm. Text tools aren't designed to handle duplicate structured pairs, so that step has to come first!

内容的提问来源于stack exchange,提问作者Jzk45

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:35:17