You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF解析场景:对含条件表达式的函数使用Mapply的技术问询

Using mapply with Conditional Expressions for PDF Data Extraction

Hey there! Let’s walk through how to use mapply with conditional expressions for your PDF data extraction workflow—sounds like you’ve got a solid setup with that CSV parameter map, so let’s make it work seamlessly with mapply.

Step 1: Prep Your CSV Data

First, let’s get your CSV structured into a format mapply can easily work with. Let’s assume your CSV looks like this (with parameters as rows, hierarchical document types as columns, and identifiers in cells):

Parameter,type1.subtypeA,type2.subtypeB.s_subtypeX,type2.subtypeC
CustomerID,CUST-ID,Customer Number,Cust#
InvoiceDate,Inv Date,Invoice Dt,Inv. Date

Use this code to load and reshape the data:

library(tidyverse)

# Load your parameter CSV
param_df <- read_csv("your_parameter_file.csv")

# Extract core components
params <- param_df$Parameter
doc_types <- colnames(param_df)[-1]  # Grab all document type columns

# Create a full list of parameter + document type pairs
param_doc_pairs <- expand.grid(
  Parameter = params,
  DocType = doc_types,
  stringsAsFactors = FALSE
)

# Match each pair to its corresponding identifier from the CSV
param_doc_pairs$Identifier <- mapply(
  function(param, doc_type) param_df[param_df$Parameter == param, doc_type],
  param_doc_pairs$Parameter,
  param_doc_pairs$DocType
)

Step 2: Build Your Conditional Extraction Function

Next, write a function that takes a parameter name, its identifier, and the hierarchical document type, then runs conditional logic based on the document’s structure. Replace the placeholder code with your actual PDF scanning/extraction logic (using tools like tesseract or pdftools):

extract_pdf_data <- function(param_name, identifier, doc_type) {
  # Split the document type into its hierarchical layers
  doc_hierarchy <- strsplit(doc_type, "\\.")[[1]]
  main_type <- doc_hierarchy[1]
  sub_type <- ifelse(length(doc_hierarchy) >= 2, doc_hierarchy[2], NA)

  # Clean conditional logic using case_when (easier to read than nested if-else)
  result <- case_when(
    main_type == "type1" ~ paste0("Extracted ", param_name, " (using '", identifier, "') from type1 document"),
    main_type == "type2" && sub_type == "subtypeB" ~ paste0("Extracted ", param_name, " (using '", identifier, "') from type2.subtypeB document"),
    main_type == "type2" && sub_type == "subtypeC" ~ paste0("Extracted ", param_name, " (using '", identifier, "') from type2.subtypeC document"),
    TRUE ~ paste0("Fallback extraction for ", param_name, " in ", doc_type, " (identifier: ", identifier, ")")
  )

  # Replace the above with your real PDF extraction code! Example:
  # pdf_path <- file.path("pdf_storage", doc_type, "target_document.pdf")
  # scanned_text <- tesseract::ocr(pdf_path)
  # result <- stringr::str_extract(scanned_text, identifier)

  return(result)
}

Step 3: Run with mapply

Now use mapply to execute your extraction function across all parameter-document pairs. mapply will iterate over each element of your input vectors and pass them to the function in sync:

# Run the extraction across all pairs
extraction_results <- mapply(
  extract_pdf_data,
  param_name = param_doc_pairs$Parameter,
  identifier = param_doc_pairs$Identifier,
  doc_type = param_doc_pairs$DocType,
  SIMPLIFY = FALSE  # Keep results as a list (better for unstructured data)
)

# Combine pairs with results into a tidy data frame for easy analysis
final_results <- cbind(param_doc_pairs, ExtractedResult = unlist(extraction_results))

Key Tips

  • Handle Missing Identifiers: Add a check like if(is.na(identifier)) { return("No identifier found") } to skip or log cases where a parameter doesn’t have an identifier for a document type.
  • Speed Up with Parallel Processing: If you’re processing hundreds of PDFs, use parallel::mclapply (Linux/macOS) or furrr::future_map (cross-platform) instead of mapply to run tasks in parallel.
  • Deep Hierarchies: For document types with 3+ sublayers, extend the doc_hierarchy parsing (e.g., s_subtype <- doc_hierarchy[3]) and add more conditions to case_when.

内容的提问来源于stack exchange,提问作者jovianlynxdroid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:06:59