如何用R高效将PDF复制的幼猫数据转换为指定格式?
高效从PDF提取幼猫数据并转换格式的R方案
核心思路
放弃文本挖掘类工具(如tm包),改用正则表达式精准提取结构化字段,配合pdftools直接读取PDF纯文本,避开干扰符号的影响,快速定位需要的信息。
所需包安装
先安装并加载必要工具包:
install.packages(c("pdftools", "dplyr", "stringr", "openxlsx")) library(pdftools) library(dplyr) library(stringr) library(openxlsx)
步骤1:读取并清理PDF文本
# 读取PDF所有页面的文本内容 pdf_text_all <- pdf_text("2013 Cats.pdf") %>% # 合并所有页面的文本 paste(collapse = "\n") %>% # 去除多余空格、换行和无意义干扰符号 str_squish() # 将文本分割为单条动物记录(以ID格式"A+数字"作为分割标记) records <- str_split(pdf_text_all, "(?=A\\d+)") %>% unlist() %>% # 过滤分割后产生的空字符串 discard(str_length(.) == 0)
步骤2:提取并转换目标字段
创建处理单条记录的函数,逐一提取所需信息:
process_record <- function(record) { # 1. 提取ID编号 id <- str_extract(record, "A\\d+") # 2. 提取年龄并转换为天数(按每月30天计算) age_match <- str_match(record, "(\\d+)m\\s*(\\d+)d") age_days <- if (!is.na(age_match[1])) { as.numeric(age_match[2])*30 + as.numeric(age_match[3]) } else { NA # 处理无年龄信息的异常记录 } # 3. 提取收容日期,拆分日/月/年 date_match <- str_match(record, "(\\d+)/(\\d+)/(\\d+)") day_intake <- as.numeric(date_match[2]) month_intake <- as.numeric(date_match[3]) year_intake <- as.numeric(date_match[4]) # 4. 提取性别(匹配M/F) sex <- str_extract(record, "\\b[MF]\\b") # 5. 提取收容类型(匹配时间后的目标词汇) intake_type <- str_extract(record, "\\d+:\\d+\\s*(?:AM|PM)\\s*(\\w+)") %>% str_remove(".*\\s") # 剔除前面的时间部分 # 返回整理后的单条记录数据框 tibble( ID_Number = id, Age_in_days = age_days, Day_of_intake = day_intake, Month_of_intake = month_intake, Year_of_intake = year_intake, Sex = sex, Intake_type = intake_type ) } # 批量处理所有记录,合并为完整数据集 cat_data <- map_dfr(records, process_record)
步骤3:导出到Excel
write.xlsx(cat_data, "幼猫收容数据汇总.xlsx", rowNames = FALSE)
微调说明
- 如果收容类型是多词格式(如
Public Drop Off),可修改intake_type的正则为"\\d+:\\d+\\s*(?:AM|PM)\\s*([A-Za-z\\s]+?)\\s*(?:Yes|No)",匹配到下一个干扰词前的完整内容。 - 若PDF存在特殊符号干扰,可在提取字段前添加
str_remove_all(record, "[^A-Za-z0-9/\\s:]")过滤非必要符号。
内容的提问来源于stack exchange,提问作者Bentley Arnold
相关产品推荐
相关产品推荐

