You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中提取含指定文本的多张表格并合并?

提取并合并符合条件的表格

步骤1:加载依赖包

先安装并加载常用数据处理工具包,如需去重完整表格还需digest包:

install.packages(c("tidyverse", "digest"))
library(tidyverse)
library(digest)

步骤2:筛选包含指定文本的表格

用purrr::keep遍历表格列表,判断每个表格的第一列是否存在my_text中的内容:

# 筛选第一列匹配my_text的表格
filtered_tabs <- keep(tabs, ~ any(.x[[1]] %in% my_text))

其中.x指代列表中的单个表格,.x[[1]]取表格第一列,any(.x[[1]] %in% my_text)用于判断该列是否有元素命中my_text中的值。

步骤3:合并筛选后的表格

用dplyr::bind_rows合并所有符合条件的表格,逻辑和Python中的pd.concat一致:

combined_df <- bind_rows(filtered_tabs)

步骤4:去重处理

根据需求选择两种去重方式:

  • 合并后去重重复行:只需合并后的数据框无重复行,用distinct():
combined_df <- combined_df %>% distinct()
  • 去重完全重复的表格:如果要先过滤列表中内容完全一致的表格,通过哈希值识别唯一表格:
# 给每个表格生成唯一哈希标识
tab_hashes <- map_chr(filtered_tabs, ~ digest(.x))
# 保留列表中唯一的表格
unique_tabs <- filtered_tabs[!duplicated(tab_hashes)]
# 合并唯一表格
combined_unique_df <- bind_rows(unique_tabs)

完整示例运行

把所有步骤整合到你的示例数据中:

my_text <- c("cash investments", "debt investments")
tabs <- NULL

tabs[[1]] <- tibble(x = sample(c("investments", "trash")), y = 1)
tabs[[2]] <- tibble(x = sample(c(my_text, "trash")), y = 2)
tabs[[3]] <- tibble(x = sample(c(my_text)), y = 3)

# 筛选+合并+去重行
filtered_tabs <- keep(tabs, ~ any(.x[[1]] %in% my_text))
combined_df <- bind_rows(filtered_tabs) %>% distinct()

print(combined_df)

运行后会自动过滤第1张不匹配的表格,输出第2、3张表格合并后的结果。

内容的提问来源于stack exchange,提问作者Raj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 19:00:06