如何修改Europe PMC文献表格获取函数以跳过无有效表格的论文并解决命名长度不匹配错误
Let's tackle your issue step by step. The error you're seeing happens because for some articles (like PMC7806552), the number of table captions retrieved by pmc_caption() doesn't match the number of tables extracted by pmc_table(). We need to add safeguards to handle this mismatch, plus make the function more robust to skip failing entries and keep processing the rest.
Modified Robust Batch Function
Here's an updated version of your function that handles the naming mismatch and other potential errors gracefully:
library(europepmc) library(tidypmc) library(tidyverse) extract_pmc_tables <- function(pmc_id) { message("-- Processing ", pmc_id, "...") # Safely retrieve the XML document doc <- tryCatch( pmc_xml(pmc_id), error = function(e) { message("------ Failed to fetch XML for ", pmc_id) return(NULL) } ) if (is.null(doc)) return(NULL) # Safely extract tables tables <- tryCatch( pmc_table(doc), error = function(e) { message("------ Failed to extract tables for ", pmc_id) return(NULL) } ) if (is.null(tables) || length(tables) == 0) { message("------ No tables found for ", pmc_id) return(NULL) } # Safely extract and match captions table_caps <- tryCatch( pmc_caption(doc) %>% filter(tag == "table"), error = function(e) { message("------ Failed to extract captions for ", pmc_id) return(NULL) } ) # Only set names if caption count matches table count if (!is.null(table_caps) && nrow(table_caps) == length(tables)) { names(tables) <- paste(table_caps$label, table_caps$text, sep = " - ") } else { # Fallback to default names if mismatch or no captions message("------ Caption-table mismatch for ", pmc_id, ", using default names") names(tables) <- paste0("Table_", seq_along(tables)) } return(tables) } # Example usage b <- epmc_search(query = 'cytarabine aml OPEN_ACCESS:Y', limit = 6) pmcids <- b$pmcid[b$isOpenAccess == "Y"] pub_tables <- lapply(pmcids, extract_pmc_tables) names(pub_tables) <- pmcids # Remove NULL entries if needed pub_tables <- pub_tables[!sapply(pub_tables, is.null)]
Key Improvements:
- Granular error handling: Each step (fetching XML, extracting tables, extracting captions) is wrapped in
tryCatchto catch specific failures and log a helpful message. - Caption-table mismatch handling: Checks if the number of captions matches the number of tables before setting names. If not, uses default numbered names instead of throwing an error.
- Clear messaging: Provides detailed feedback about what's happening with each PMC ID, making it easier to debug issues.
- Graceful skips: Returns
NULLfor any failing entry instead of stopping the entire batch process.
Simplified Alternative with purrr
If you prefer a more concise approach using purrr's safe mapping functions, here's a streamlined version:
library(purrr) # Create a safely wrapped version of the extraction function safe_extract <- safely(extract_pmc_tables) # Process all PMC IDs in one go pub_tables_safe <- map(pmcids, safe_extract) # Separate results and errors results <- map(pub_tables_safe, "result") %>% compact() # Keep only non-NULL results errors <- map(pub_tables_safe, "error") %>% discard(is.null) # Get only error details # Name the results names(results) <- pmcids[!sapply(results, is.null)]
This uses safely() to capture both results and errors without stopping execution. You can then inspect the errors list to see exactly which PMC IDs failed and why.
Why Your Original Error Happened
For PMC7806552, either:
- Some tables don't have associated captions, or
pmc_caption()is picking up extra entries (like supplementary tables or non-standard captions) that don't correspond to the tables extracted bypmc_table().
The modified function handles both cases by falling back to default names instead of crashing.
内容的提问来源于stack exchange,提问作者PesKchan

