You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中基于两个dataframe生成指定词集的文档词频矩阵

R语言实现自定义目标词的文档词频矩阵

以下给出两种可直接运行的实现方案:

方案1:tidyverse生态实现(代码简洁易读)

第一步:构造示例测试数据

df1 <- data.frame(
  category = c("person1", "person2", "person3"),
  text = c("hello word I like turtles", "re: turtles! I think turtles are stellar!", "sunflowers are nice.")
)

df2 <- data.frame(
  col1 = c("x", "y", "w", "f"),
  term = c("turtles", "hello", "sunflowers", "I")
)

第二步:核心统计代码

# 加载依赖包
library(dplyr)
library(stringr)

target_terms <- df2$term
result <- df1 %>%
  # 按行处理每个主体的文本
  rowwise() %>%
  # 遍历所有目标词统计出现次数,添加单词边界避免部分匹配
  mutate(across(all_of(target_terms), ~ str_count(text, paste0("\\b", .x, "\\b")))) %>%
  # 移除原始文本列,得到最终结果
  select(-text) %>%
  ungroup()

方案2:base R实现(无需安装额外依赖包)

target_terms <- df2$term
# 初始化计数矩阵
count_matrix <- matrix(
  0, 
  nrow = nrow(df1), 
  ncol = length(target_terms),
  dimnames = list(df1$category, target_terms)
)

# 遍历每个目标词统计频次
for (i in seq_along(target_terms)) {
  pattern <- paste0("\\b", target_terms[i], "\\b")
  count_matrix[, i] <- sapply(gregexpr(pattern, df1$text), function(match_res) sum(match_res > 0))
}

# 转换为要求的dataframe格式
result <- cbind(data.frame(category = rownames(count_matrix)), count_matrix)
rownames(result) <- NULL

两种方案运行后得到的result和你给出的预期输出结构完全一致。

内容的提问来源于stack exchange,提问作者mk2080

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 02:15:02