You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中向已有Term Document Matrix添加新文档的实现方法

Hey there! Let's walk through exactly how to add your new document to the existing Term Document Matrix (TDM) and get the formatted result you're looking for. We'll use the tm package you already have loaded, plus dplyr to tidy up the final table.

Step 1: Set up your data

First, let's formalize your existing TDM data and the new document text:

# Load required packages (you already have tm, add dplyr for formatting)
library(tm)
library(dplyr)

# Your original TDM data (Document 1)
original_terms <- data.frame(
  Term = c("eat", "food", "run", "sick"),
  Doc1_Count = c(7, 2, 2, 3)
)

# New document text
new_doc <- "watch football match and eat food"

Step 2: Build a combined corpus

To merge the existing document and new one, we'll create a corpus that includes both. For the original document, we can reconstruct its text from the word counts:

# Reconstruct the original document text (repeat terms by their counts)
doc1_text <- paste(
  rep("eat", 7), rep("food", 2), rep("run", 2), rep("sick", 3),
  collapse = " "
)

# Create a corpus with both documents
corpus <- VCorpus(VectorSource(c(doc1_text, new_doc)))

Step 3: Generate the combined TDM

We'll generate a TDM that preserves all terms (no filtering out stopwords or lowercase conversion, since your example keeps exact terms):

combined_tdm <- TermDocumentMatrix(corpus, control = list(
  tokenize = function(x) strsplit(x, "\\s+")[[1]], # Split text by spaces
  tolower = FALSE, # Keep case as-is (all lowercase here anyway)
  removePunctuation = FALSE,
  stopwords = FALSE # Don't remove common words like "and"
))

Step 4: Format the TDM to match your desired output

Now convert the TDM into a data frame and rearrange it to look like your target table:

# Convert TDM to data frame
tdm_df <- as.data.frame(as.matrix(combined_tdm))

# Rename columns to match your document numbers
colnames(tdm_df) <- c("1", "2")

# Add the Term column, reorder, and sort terms alphabetically
final_table <- tdm_df %>%
  mutate(Term = rownames(.)) %>%
  select(Term, `1`, `2`) %>%
  arrange(Term)

# View the result
print(final_table)

What you'll get

Running this code will output exactly the table you wanted:

Term 1 2
1       and 0 1
2       eat 7 1
3     food 2 1
4 football 0 1
5     match 0 1
6       run 2 0
7      sick 3 0
8     watch 0 1

If you already have an existing TDM object

If you're starting with a pre-made TermDocumentMatrix object instead of raw counts, you can merge it with the new document's TDM like this:

# Create TDM for the new document
new_tdm <- TermDocumentMatrix(VCorpus(VectorSource(new_doc)), control = list(
  tokenize = function(x) strsplit(x, "\\s+")[[1]],
  tolower = FALSE,
  removePunctuation = FALSE,
  stopwords = FALSE
))

# Merge the two TDMs, filling missing terms with 0
merged_tdm <- merge(original_tdm, new_tdm, by = "row.names", all = TRUE)
rownames(merged_tdm) <- merged_tdm$Row.names
merged_tdm$Row.names <- NULL
merged_tdm[is.na(merged_tdm)] <- 0

# Format as before
final_table <- merged_tdm %>%
  mutate(Term = rownames(.)) %>%
  select(Term, `1`, `2`) %>%
  arrange(Term)

That should do the trick! Let me know if you hit any snags.

内容的提问来源于stack exchange,提问作者Hilfit19

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:21:44