PDF解析场景:对含条件表达式的函数使用Mapply的技术问询
mapply with Conditional Expressions for PDF Data Extraction Hey there! Let’s walk through how to use mapply with conditional expressions for your PDF data extraction workflow—sounds like you’ve got a solid setup with that CSV parameter map, so let’s make it work seamlessly with mapply.
Step 1: Prep Your CSV Data
First, let’s get your CSV structured into a format mapply can easily work with. Let’s assume your CSV looks like this (with parameters as rows, hierarchical document types as columns, and identifiers in cells):
Parameter,type1.subtypeA,type2.subtypeB.s_subtypeX,type2.subtypeC CustomerID,CUST-ID,Customer Number,Cust# InvoiceDate,Inv Date,Invoice Dt,Inv. Date
Use this code to load and reshape the data:
library(tidyverse) # Load your parameter CSV param_df <- read_csv("your_parameter_file.csv") # Extract core components params <- param_df$Parameter doc_types <- colnames(param_df)[-1] # Grab all document type columns # Create a full list of parameter + document type pairs param_doc_pairs <- expand.grid( Parameter = params, DocType = doc_types, stringsAsFactors = FALSE ) # Match each pair to its corresponding identifier from the CSV param_doc_pairs$Identifier <- mapply( function(param, doc_type) param_df[param_df$Parameter == param, doc_type], param_doc_pairs$Parameter, param_doc_pairs$DocType )
Step 2: Build Your Conditional Extraction Function
Next, write a function that takes a parameter name, its identifier, and the hierarchical document type, then runs conditional logic based on the document’s structure. Replace the placeholder code with your actual PDF scanning/extraction logic (using tools like tesseract or pdftools):
extract_pdf_data <- function(param_name, identifier, doc_type) { # Split the document type into its hierarchical layers doc_hierarchy <- strsplit(doc_type, "\\.")[[1]] main_type <- doc_hierarchy[1] sub_type <- ifelse(length(doc_hierarchy) >= 2, doc_hierarchy[2], NA) # Clean conditional logic using case_when (easier to read than nested if-else) result <- case_when( main_type == "type1" ~ paste0("Extracted ", param_name, " (using '", identifier, "') from type1 document"), main_type == "type2" && sub_type == "subtypeB" ~ paste0("Extracted ", param_name, " (using '", identifier, "') from type2.subtypeB document"), main_type == "type2" && sub_type == "subtypeC" ~ paste0("Extracted ", param_name, " (using '", identifier, "') from type2.subtypeC document"), TRUE ~ paste0("Fallback extraction for ", param_name, " in ", doc_type, " (identifier: ", identifier, ")") ) # Replace the above with your real PDF extraction code! Example: # pdf_path <- file.path("pdf_storage", doc_type, "target_document.pdf") # scanned_text <- tesseract::ocr(pdf_path) # result <- stringr::str_extract(scanned_text, identifier) return(result) }
Step 3: Run with mapply
Now use mapply to execute your extraction function across all parameter-document pairs. mapply will iterate over each element of your input vectors and pass them to the function in sync:
# Run the extraction across all pairs extraction_results <- mapply( extract_pdf_data, param_name = param_doc_pairs$Parameter, identifier = param_doc_pairs$Identifier, doc_type = param_doc_pairs$DocType, SIMPLIFY = FALSE # Keep results as a list (better for unstructured data) ) # Combine pairs with results into a tidy data frame for easy analysis final_results <- cbind(param_doc_pairs, ExtractedResult = unlist(extraction_results))
Key Tips
- Handle Missing Identifiers: Add a check like
if(is.na(identifier)) { return("No identifier found") }to skip or log cases where a parameter doesn’t have an identifier for a document type. - Speed Up with Parallel Processing: If you’re processing hundreds of PDFs, use
parallel::mclapply(Linux/macOS) orfurrr::future_map(cross-platform) instead ofmapplyto run tasks in parallel. - Deep Hierarchies: For document types with 3+ sublayers, extend the
doc_hierarchyparsing (e.g.,s_subtype <- doc_hierarchy[3]) and add more conditions tocase_when.
内容的提问来源于stack exchange,提问作者jovianlynxdroid

