You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让tabulapdf以字符类型读取PDF表格的Quantity列

解决tabulapdf提取PDF表格时数值末尾零丢失的问题

可以通过指定col_types参数强制将Quantity列以字符类型读取,避免自动类型转换导致的末尾零丢失。修改后的代码如下:

url <- "https://www.fenabrave.org.br/portal/files/2023_01_2.pdf"

download.file(url, 'cars.pdf', mode="wb")

library(tabulapdf)
library(dplyr)

df <- extract_tables(
  'cars.pdf',
  pages = 27,
  area = list(c(126.4826, 96.5997, 782.1684, 297.9600)),
  guess = FALSE,
  col_types = cols(
    Model = col_character(),
    Quantity = col_character()
  )
) %>% 
  bind_rows() %>% 
  set_names(c("Model","Quantity"))

说明

  • col_types参数可明确指定每一列的数据类型,这里将Quantity列设为字符型,能完整保留原始文本中的末尾零(比如"2.830"、"1.220"),不会被自动转换为数值类型而丢失信息。
  • 后续如果需要对Quantity列做数值计算,可再根据需求用as.numeric()或parse_number()等函数手动转换类型。

内容的提问来源于stack exchange,提问作者AndreASousa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 21:32:40