R语言提取PDF表格:现有代码繁琐,求更高效实现方案
Hey there! Let's turn that repetitive, verbose code into something cleaner and more maintainable. Your current approach manually slices and converts segments of the table, which gets tedious fast—here are a couple of streamlined methods to handle this better.
First: Use Tidyverse Tools for Batch Processing
The core issue with your code is duplicated logic for each table segment. We can fix this by defining your row ranges once, then using purrr to process all segments in one go, and dplyr to merge the results automatically.
First, make sure you have the necessary packages installed and loaded:
install.packages(c("tabulizer", "dplyr", "purrr")) library(tabulizer) library(dplyr) library(purrr)
Then rewrite your extraction workflow like this:
# Define your PDF source pdf_url <- "http://www.drustvo-antropologov.si/AN/PDF/2012_2/Anthropological_Notebooks_XVIII_2_Bjelica.pdf" # Extract and format the full table once full_table <- extract_tables(pdf_url) %>% unlist() %>% matrix(nrow = 195, byrow = TRUE) %>% as.data.frame(stringsAsFactors = FALSE) # Avoid factor columns by default # List all the row ranges you need to extract row_segments <- list( c(start = 29, end = 52), c(start = 53, end = 76), c(start = 101, end = 126), c(start = 127, end = 152) ) # Process all segments and combine into one data frame combined_table <- row_segments %>% map(~ full_table[.x["start"]:.x["end"], ]) %>% bind_rows()
Why this is better:
- No repeated code: You only write the logic for slicing once, even if you add more segments later.
- Readability: Anyone looking at the code can immediately see which row ranges you're targeting.
- Easier maintenance: If you need to adjust a row number or add a new segment, just update the
row_segmentslist.
Bonus: Add Source Context (Optional)
If you want to track which original segment each row came from, use imap_dfr to add an identifier column:
combined_table_with_source <- row_segments %>% imap_dfr(function(segment, segment_id) { full_table[segment["start"]:segment["end"], ] %>% mutate(source_segment = paste0("Segment_", segment_id)) })
Check if You Can Avoid Manual Row Slicing
One thing to double-check: does extract_tables() already return the individual sub-tables as separate elements in the list? Sometimes tabulizer can detect distinct tables in the PDF automatically, which would let you skip the manual row slicing entirely:
# Try extracting tables directly as separate elements auto_detected_tables <- extract_tables(pdf_url) # If this works, combine them directly combined_table <- auto_detected_tables %>% map(~ as.data.frame(.x, stringsAsFactors = FALSE)) %>% bind_rows()
This would be the most elegant solution if tabulizer can recognize the sub-tables on its own. If not, the batch processing method above is your next best bet.
内容的提问来源于stack exchange,提问作者Helena

