You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分R语言提取PDF表格中合并的多列并设置对应列名

拆分PDF提取表格的合并列并重命名

我用R的extract_tables函数提取《UBI Kenya paper.pdf》第59-68页的表格时,调用as.data.frame报错,改用data.table::as.data.table转换格式。截取table_E.9的行数据后,发现V2-V5列中,零售贸易、制造业、交通运输、服务业这几个类别的“企业数量(# Enterprises)”与“净收入(Net Revenues)”值合并在同一个单元格里。需要将第4-6行(对应Long Term Arm、其标准误、Short Term Arm)的V2-V5列各拆分为两列,并设置列名为“零售贸易 - 企业数量”“零售贸易 - 净收入”等对应名称,还原表格原始结构。

以下是table_E.9的dput结果:

structure(list(V1 = c("", "", "", "Long Term Arm", "", "Short Term Arm"), 
               V2 = c("Retail Trade", "# Enterprises Net Revenues", "(1) (2)", "3.89*** 1601.42*", "[1.28] [824.74]", "2.34** 464.60*"), 
               V3 = c("Manufacturing", "# Enterprises Net Revenues", "(3) (4)", "0.02 51.90", "[.27] [120.79]", "0.03 17.82"), 
               V4 = c("Transportation", "# Enterprises Net Revenues", "(5) (6)", "0.53* 100.76", "[.29] [85.37]", "-0.12 -3.85"), 
               V5 = c("Services", "# Enterprises Net Revenues", "(7) (8)", "0.23 198.64", "[.33] [205.6]", "-0.04 70.95")), 
          row.names = c(NA, -6L), class = c("data.table", "data.frame"), .internal.selfref = <pointer: 0x000001716fd35930>)

处理步骤

  1. 提取列名信息:从V2-V5的第一行获取行业名称,第二行获取指标名称,组合成新列名
  2. 拆分合并的单元格:针对第4-6行的V2-V5列,按空格拆分每个单元格内容为两个值
  3. 重组表格结构:将拆分后的数据与V1列整合,设置新列名
library(data.table)

# 加载数据
dt <- structure(list(V1 = c("", "", "", "Long Term Arm", "", "Short Term Arm"), 
               V2 = c("Retail Trade", "# Enterprises Net Revenues", "(1) (2)", "3.89*** 1601.42*", "[1.28] [824.74]", "2.34** 464.60*"), 
               V3 = c("Manufacturing", "# Enterprises Net Revenues", "(3) (4)", "0.02 51.90", "[.27] [120.79]", "0.03 17.82"), 
               V4 = c("Transportation", "# Enterprises Net Revenues", "(5) (6)", "0.53* 100.76", "[.29] [85.37]", "-0.12 -3.85"), 
               V5 = c("Services", "# Enterprises Net Revenues", "(7) (8)", "0.23 198.64", "[.33] [205.6]", "-0.04 70.95")), 
          row.names = c(NA, -6L), class = c("data.table", "data.frame"))

# 生成新列名:行业名称 + 指标名称
sector_names <- unlist(dt[1, c(V2, V3, V4, V5)])
metric_names <- strsplit(dt[2, V2], " ")[[1]]
new_colnames <- apply(expand.grid(sector_names, metric_names), 1, function(x) paste(x[1], x[2], sep = " - "))

# 处理需要拆分的行:第4到6行
split_rows <- dt[4:6, ]
# 拆分V2-V5列的合并内容
split_data <- lapply(split_rows[, -1], function(col) unlist(strsplit(col, " +"))) |> 
  do.call(cbind, args = _)
colnames(split_data) <- new_colnames

# 整合V1列与拆分后的数据
result <- cbind(split_rows[, .(V1)], as.data.table(split_data))

# 查看最终结果
print(result)

结果说明

处理后的表格保留了V1列的分组标识,每个行业的“企业数量”和“净收入”各占独立列,列名清晰对应原始表格的结构,方便后续分析使用。

内容的提问来源于stack exchange,提问作者hks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 04:52:07