You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量提取Word/PDF抗体数据单文本并转换为R数据框或Excel表格

Batch Extract Antibody Datasheet Fields to DataFrame/Excel in R

Overview

Since your 150 Word and PDF datasheets have consistent formatting, we can leverage regex and R's document-processing packages to pull the exact fields you need (Catalogue No., Name, Source) into a structured dataframe, then export to Excel if needed. Let's break this down into manageable steps.

Step 1: Install & Load Required Packages

First, get all the tools we'll need for PDFs, Word docs, data manipulation, and export:

install.packages(c("pdftools", "officer", "docxtractr", "dplyr", "writexl", "stringr"))
library(pdftools)
library(officer)
library(docxtractr)
library(dplyr)
library(writexl)
library(stringr)

Step 2: Process PDF Files

You already started with pdftools, so we'll build on that. We'll create a reusable function to extract fields from a single PDF, then loop through all PDFs in your folder.

Function to Extract PDF Fields

extract_pdf_fields <- function(pdf_path) {
  # Pull full text from the PDF and collapse into one string
  pdf_text <- pdf_text(pdf_path) %>% paste(collapse = "\n")
  
  # Use regex to target each field (adjust patterns if your labels vary slightly)
  catalogue_no <- str_extract(pdf_text, regex("Catalogue No\\.:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>%
    str_remove("Catalogue No\\.:\\s*") %>% str_trim()
  
  name <- str_extract(pdf_text, regex("Name:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>%
    str_remove("Name:\\s*") %>% str_trim()
  
  source <- str_extract(pdf_text, regex("Source:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>%
    str_remove("Source:\\s*") %>% str_trim()
  
  # Return a single row of data
  tibble(
    File_Type = "PDF",
    File_Name = basename(pdf_path),
    Catalogue_No = catalogue_no,
    Name = name,
    Source = source
  )
}

Batch Process All PDFs

# Replace with your actual folder path
folder_path <- "path/to/your/datasheets"

# Get all PDF files in the folder
pdf_files <- list.files(path = folder_path, pattern = "\\.pdf$", full.names = TRUE)

# Apply the function to every PDF and combine into a dataframe
pdf_data <- lapply(pdf_files, extract_pdf_fields) %>% bind_rows()

Step 3: Process Word (.docx) Files

For Word documents, we'll use docxtractr to extract text, then apply the same regex logic as PDFs.

Function to Extract Word Fields

extract_word_fields <- function(docx_path) {
  # Pull full text from the Word doc
  docx_text <- docx_extract_all(docx_path)$text %>% paste(collapse = "\n")
  
  # Reuse the same regex patterns as PDFs (consistent formatting = less work!)
  catalogue_no <- str_extract(docx_text, regex("Catalogue No\\.:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>%
    str_remove("Catalogue No\\.:\\s*") %>% str_trim()
  
  name <- str_extract(docx_text, regex("Name:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>%
    str_remove("Name:\\s*") %>% str_trim()
  
  source <- str_extract(docx_text, regex("Source:\\s*(.*?)(\\n|\\s{2,})", dotall = TRUE)) %>%
    str_remove("Source:\\s*") %>% str_trim()
  
  # Return a single row of data
  tibble(
    File_Type = "Word",
    File_Name = basename(docx_path),
    Catalogue_No = catalogue_no,
    Name = name,
    Source = source
  )
}

Batch Process All Word Files

# Get all .docx files in the folder
word_files <- list.files(path = folder_path, pattern = "\\.docx$", full.names = TRUE)

# Apply the function to every Word doc and combine into a dataframe
word_data <- lapply(word_files, extract_word_fields) %>% bind_rows()

Step 4: Combine & Clean the Data

Merge the PDF and Word datasets into one, then handle any missing values or minor inconsistencies:

# Combine both datasets
all_data <- bind_rows(pdf_data, word_data)

# Replace NA values with a clear label (adjust based on your needs)
all_data <- all_data %>%
  mutate(across(c(Catalogue_No, Name, Source), ~ifelse(is.na(.), "Field not found", .)))

Step 5: Export to DataFrame or Excel

Keep as R DataFrame

The all_data object is already a tibble (a modern R dataframe) that you can use for further analysis in R.

Export to Excel

# Save to an Excel file in your working directory
write_xlsx(all_data, path = "antibody_datasheet_summary.xlsx")

Quick Notes for Edge Cases

  • If some datasheets use slight variations of field labels (e.g., "Catalog No." instead of "Catalogue No."), update the regex to include both: regex("Catalogue No\\.|Catalog No\\.:\\s*(.*?)(\\n|\\s{2,})")
  • If fields span multiple lines, the dotall = TRUE flag in the regex already handles this by allowing . to match newlines.
  • For older .doc files (not .docx), convert them to .docx first (most word processors can do this), or use the tm package with antiword support.

内容的提问来源于stack exchange,提问作者Filipe Rigueiro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:21:15