网页抓取编译数据时缺失值处理及批量代码优化咨询
Hey there! Let's tackle your two web scraping challenges step by step—handling missing values and streamlining batch processing. Here's a robust, efficient solution using R's rvest package:
The core problem is that when html_nodes() can't locate an element (e.g., no email link on page 3), html_text() returns an empty vector. This breaks cbind() because all columns need matching lengths.
To fix this, we'll add a safety check for each field: if the element doesn't exist, we'll return NA instead of an empty value. This keeps all columns consistent in length, so your data compilation won't throw errors.
Instead of writing repetitive code for each page, we'll restructure your workflow to:
- Store all your DOI URLs in a single vector
- Create a reusable function to scrape one page and return a structured data frame row
- Use batch processing tools to run this function on every URL automatically
Complete Improved Code
library(rvest) library(purrr) # For streamlined batch processing # Step 1: List all your DOI URLs here doi_urls <- c( "http://journals.sagepub.com/doi/full/10.3102/0013189X08317501", "http://journals.sagepub.com/doi/abs/10.3102/0013189X037002102", "http://journals.sagepub.com/doi/abs/10.3102/0013189X037002104", # Add more URLs in this format as needed "http://journals.sagepub.com/doi/full/xxxx/xxxxxxx" ) # Step 2: Build a function to scrape a single journal page scrape_single_page <- function(url) { # Read the page (add a small delay if needed to avoid overloading the server) page <- read_html(url) # Helper function to safely extract text (returns NA if element is missing) safe_extract <- function(selector, xpath = NULL) { element <- if (!is.null(xpath)) { html_nodes(page, xpath = xpath) } else { html_nodes(page, selector) } if (length(element) == 0) { NA_character_ } else { html_text(element, trim = TRUE) } } # Extract all fields and return as a data frame row data.frame( author = safe_extract('.art_authors'), year = safe_extract('.year'), journalName = safe_extract('.journalName'), art_title = safe_extract('.art_title'), volume = safe_extract('.volume'), page = safe_extract('.page'), email = safe_extract(xpath = "//a[@class = 'email']"), source_url = url, # Optional: keep the URL for reference stringsAsFactors = FALSE ) } # Step 3: Batch process all URLs and combine into one dataset final_dataset <- map_df(doi_urls, scrape_single_page) # View the compiled results View(final_dataset)
Key Improvements Explained
- Missing Value Safety: The
safe_extract()function ensures every field returns a value (evenNAif the element is missing), so your data frame never breaks during compilation. - Batch Efficiency: No more copying/pasting code for each page—just add new URLs to the
doi_urlsvector, and the function handles the rest. - Maintainable Code: Using a function makes it easy to add new fields (e.g., abstract, keywords) later by updating the
data.frame()call. - Polite Scraping: You can add
Sys.sleep(1)inside thescrape_single_pagefunction to pause between requests, which helps avoid getting blocked by the journal's server.
内容的提问来源于stack exchange,提问作者JWH2006

