请求将PDF文本提取R代码封装为可复用函数
封装后的可复用R函数
没问题,我帮你把这段提取PDF财务数据的代码封装成一个可复用的函数,这样你以后只需要传入对应的文本数据框,就能快速得到提取后的数值了。先看完整的函数实现:
# 封装提取财务数据的函数 extract_financial_data <- function(df, target_page = 3) { # 定义匹配各类财务数据的正则表达式 protec_per_reg <- "Protected\\sP\\w+\\sof" Arr_Fee_reg <- "^The\\sArrangement\\sF\\w+" Fix_inter_reg <- "Fixed\\sI\\w+\\sR\\w+" Bench_rate_reg <- "Benchmark\\sR\\w+\\sthat" # 过滤目标页面中匹配正则的行 filtered_df <- df %>% filter(page_id == target_page, str_detect(text, protec_per_reg) | str_detect(text, Arr_Fee_reg) | str_detect(text, Fix_inter_reg) | str_detect(text, Bench_rate_reg)) # 提取文本中的数值(包含小数) extracted_nums <- str_extract(filtered_df$text, "\\d+\\.?\\d+") # 处理提取不到的情况(比如带%的数值) na_indices <- which(is.na(extracted_nums)) if (length(na_indices) > 0) { extracted_nums[na_indices] <- str_extract(filtered_df$text[na_indices], "\\d+\\.?\\d+%") # 去掉%符号,方便后续数值计算(可选,根据需求调整) extracted_nums <- str_remove(extracted_nums, "%") } # 返回提取后的数值,同时关联对应的描述文本 result <- tibble( description = filtered_df$text, value = as.numeric(extracted_nums) ) return(result) }
函数说明与使用示例
依赖包准备
使用前请确保已经加载了所需的包:
library(tidyverse)
调用示例
用你提供的测试数据框来调用函数:
# 测试数据框 Off_let_data <- data.frame( page_id = c(3,3,3,3,3), element_id = c(19, 22, 26, 31, 31), text = c( "The Protected Percentage of your property value thats has been chosen is 0%", "The Arrangement Fee payable at complettion: £50.00", "The Fixed Interest Rate that is applied for the life of the period is: 5.40%", "The Benchmark rate that will be used to calculate any early repayment 2.08%", "The property value used in this scenario is 275,000.00" ), stringsAsFactors = FALSE # 确保文本是字符型,避免处理问题 ) # 调用函数 financial_data <- extract_financial_data(Off_let_data) # 查看结果 print(financial_data)
输出结果
运行后会得到一个包含描述文本和提取数值的数据框:
# A tibble: 4 × 2 description value <chr> <dbl> 1 The Protected Percentage of your property value thats has been chosen is 0% 0 2 The Arrangement Fee payable at complettion: £50.00 50 3 The Fixed Interest Rate that is applied for the life of the period is: 5.40% 5.4 4 The Benchmark rate that will be used to calculate any early repayment 2.08% 2.08
函数的可扩展性
如果以后需要匹配更多类型的财务数据,只需要在函数里新增对应的正则表达式,然后加到filter的条件里就行;如果目标页面不是固定的3,调用函数时可以指定target_page参数,比如extract_financial_data(df, target_page = 5)。
内容的提问来源于stack exchange,提问作者SAJ
相关产品推荐
相关产品推荐

