如何在html_nodes循环中处理无效URL,填充空值并继续执行?
Hey there! I totally get the frustration—nothing kills productivity faster than a loop grinding to a halt halfway through 200k URLs because of one bad link. Let’s fix this with error handling so your scraper keeps chugging along, even when it hits unreachable addresses.
The Core Issue
Your current code crashes because read_html() throws an error when it can’t resolve a host or access a URL, and unhandled errors break the loop entirely. We need to catch those errors and tell R to keep going instead.
Fixed Code with Error Handling
Here’s a revised version of your loop using tryCatch() to handle failures gracefully, plus a fix for storing results (your original code was overwriting a each time, so you’d only end up with the last URL’s data):
library(rvest) # Initialize a list to store results (way more efficient for 200k entries) results <- vector("list", length(url)) for (i in seq_along(url)) { current_url <- url[i] # Use tryCatch to capture errors and continue extracted_text <- tryCatch({ # Normal execution: scrape the URL page <- read_html(current_url) page %>% html_nodes("p") %>% html_text() }, error = function(e) { # If something goes wrong, log the issue and return empty text message(paste("Skipping invalid URL:", current_url, "\nError:", e$message)) character(0) # Returns an empty character vector; use "" if you want a single empty string }) # Store the result for this URL results[[i]] <- extracted_text } # Optional: If you want to combine all results into a single vector # combined_text <- unlist(results)
Key Improvements
- Error Catching:
tryCatch()lets us define exactly what happens when an error occurs—here, we log the problematic URL and return an empty value instead of crashing. - Proper Storage: Using a pre-allocated list (
vector("list", length(url))) is way faster than dynamically growing a list, which is critical for 200k entries. - Transparency: The
message()call lets you keep track of which URLs failed, so you can revisit them later if needed.
Bonus: Faster Parallel Processing (For 200k URLs)
A single loop will take forever with 200k URLs. For a speed boost, use the furrr package to run scrapes in parallel, paired with possibly() (a simpler alternative to tryCatch() for functional code):
library(rvest) library(furrr) # Set up parallel sessions (adjust based on your CPU cores) plan(multisession, workers = 4) # Define a "safe" version of your scraping function that returns empty on failure safe_scrape_p <- possibly(function(link) { read_html(link) %>% html_nodes("p") %>% html_text() }, otherwise = character(0)) # Run parallel scraping with a progress bar results <- future_map(url, safe_scrape_p, .progress = TRUE)
This will cut down your runtime drastically—just make sure to add a small delay (like Sys.sleep(0.1)) inside the scraping function if you’re hitting a single domain, to avoid getting blocked by anti-scraping measures.
内容的提问来源于stack exchange,提问作者CatCaller

