You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R的PDE包从PDF中提取指定标题的单张表格?

仅提取指定PDF表格的解决方案

可以只提取目标表格,给你两种实用方案:

方案一:用PDE_pdfs2table_searchandfilter精准匹配

之前提取出大量表格大概率是因为关键词匹配太宽泛,加上精确匹配参数就能锁定目标表格:

library(PDE)
# 精准定位目标表格
target_table <- PDE_pdfs2table_searchandfilter(
  pdf = 'GPI-2023-Web.pdf',
  searchString = 'Table 1.1: Safety and Security domain',
  exact = TRUE,       # 开启精确匹配,完全匹配标题
  caseSensitive = TRUE # 保持大小写一致,避免误匹配类似标题
)
# 将提取到的表格保存为CSV
if(length(target_table) > 0) {
  write.csv(target_table[[1]], "target_table.csv", row.names = FALSE)
}

方案二:先全提再筛选(更稳妥)

如果第一种方案还是有问题,先提取所有表格及元信息,再手动筛选目标表格:

library(PDE)
# 提取所有表格,同时返回表格的元数据(含标题)
all_tables <- PDE_pdfs2table(pdf = 'GPI-2023-Web.pdf', returnInfo = TRUE)
# 找到标题完全匹配的表格索引
target_index <- which(sapply(all_tables$info, function(table_meta) {
  table_meta$tableName == 'Table 1.1: Safety and Security domain'
}))
# 提取并保存目标表格
if(length(target_index) > 0) {
  target_table <- all_tables$tables[[target_index]]
  write.csv(target_table, "target_table.csv", row.names = FALSE)
}

内容的提问来源于stack exchange,提问作者Gopala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 08:17:21