批量提取Word/PDF抗体数据单文本并转换为R数据框或Excel表格
Overview
Since your 150 Word and PDF datasheets have consistent formatting, we can leverage regex and R's document-processing packages to pull the exact fields you need (Catalogue No., Name, Source) into a structured dataframe, then export to Excel if needed. Let's break this down into manageable steps.
Step 1: Install & Load Required Packages
First, get all the tools we'll need for PDFs, Word docs, data manipulation, and export:
install.packages(c("pdftools", "officer", "docxtractr", "dplyr", "writexl", "stringr")) library(pdftools) library(officer) library(docxtractr) library(dplyr) library(writexl) library(stringr)
Step 2: Process PDF Files
You already started with pdftools, so we'll build on that. We'll create a reusable function to extract fields from a single PDF, then loop through all PDFs in your folder.
Function to Extract PDF Fields
extract_pdf_fields <- function(pdf_path) { # Pull full text from the PDF and collapse into one string pdf_text <- pdf_text(pdf_path) %>% paste(collapse = "\n") # Use regex to target each field (adjust patterns if your labels vary slightly) catalogue_no <- str_extract(pdf_text, regex("Catalogue No\\.:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>% str_remove("Catalogue No\\.:\\s*") %>% str_trim() name <- str_extract(pdf_text, regex("Name:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>% str_remove("Name:\\s*") %>% str_trim() source <- str_extract(pdf_text, regex("Source:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>% str_remove("Source:\\s*") %>% str_trim() # Return a single row of data tibble( File_Type = "PDF", File_Name = basename(pdf_path), Catalogue_No = catalogue_no, Name = name, Source = source ) }
Batch Process All PDFs
# Replace with your actual folder path folder_path <- "path/to/your/datasheets" # Get all PDF files in the folder pdf_files <- list.files(path = folder_path, pattern = "\\.pdf$", full.names = TRUE) # Apply the function to every PDF and combine into a dataframe pdf_data <- lapply(pdf_files, extract_pdf_fields) %>% bind_rows()
Step 3: Process Word (.docx) Files
For Word documents, we'll use docxtractr to extract text, then apply the same regex logic as PDFs.
Function to Extract Word Fields
extract_word_fields <- function(docx_path) { # Pull full text from the Word doc docx_text <- docx_extract_all(docx_path)$text %>% paste(collapse = "\n") # Reuse the same regex patterns as PDFs (consistent formatting = less work!) catalogue_no <- str_extract(docx_text, regex("Catalogue No\\.:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>% str_remove("Catalogue No\\.:\\s*") %>% str_trim() name <- str_extract(docx_text, regex("Name:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>% str_remove("Name:\\s*") %>% str_trim() source <- str_extract(docx_text, regex("Source:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>% str_remove("Source:\\s*") %>% str_trim() # Return a single row of data tibble( File_Type = "Word", File_Name = basename(docx_path), Catalogue_No = catalogue_no, Name = name, Source = source ) }
Batch Process All Word Files
# Get all .docx files in the folder word_files <- list.files(path = folder_path, pattern = "\\.docx$", full.names = TRUE) # Apply the function to every Word doc and combine into a dataframe word_data <- lapply(word_files, extract_word_fields) %>% bind_rows()
Step 4: Combine & Clean the Data
Merge the PDF and Word datasets into one, then handle any missing values or minor inconsistencies:
# Combine both datasets all_data <- bind_rows(pdf_data, word_data) # Replace NA values with a clear label (adjust based on your needs) all_data <- all_data %>% mutate(across(c(Catalogue_No, Name, Source), ~ifelse(is.na(.), "Field not found", .)))
Step 5: Export to DataFrame or Excel
Keep as R DataFrame
The all_data object is already a tibble (a modern R dataframe) that you can use for further analysis in R.
Export to Excel
# Save to an Excel file in your working directory write_xlsx(all_data, path = "antibody_datasheet_summary.xlsx")
Quick Notes for Edge Cases
- If some datasheets use slight variations of field labels (e.g., "Catalog No." instead of "Catalogue No."), update the regex to include both:
regex("Catalogue No\\.|Catalog No\\.:\\s*(.*?)(\\n|\\s{2,})") - If fields span multiple lines, the
dotall = TRUEflag in the regex already handles this by allowing.to match newlines. - For older
.docfiles (not.docx), convert them to.docxfirst (most word processors can do this), or use thetmpackage with antiword support.
内容的提问来源于stack exchange,提问作者Filipe Rigueiro

