You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在html_nodes循环中处理无效URL,填充空值并继续执行?

Hey there! I totally get the frustration—nothing kills productivity faster than a loop grinding to a halt halfway through 200k URLs because of one bad link. Let’s fix this with error handling so your scraper keeps chugging along, even when it hits unreachable addresses.

The Core Issue

Your current code crashes because read_html() throws an error when it can’t resolve a host or access a URL, and unhandled errors break the loop entirely. We need to catch those errors and tell R to keep going instead.

Fixed Code with Error Handling

Here’s a revised version of your loop using tryCatch() to handle failures gracefully, plus a fix for storing results (your original code was overwriting a each time, so you’d only end up with the last URL’s data):

library(rvest)

# Initialize a list to store results (way more efficient for 200k entries)
results <- vector("list", length(url))

for (i in seq_along(url)) {
  current_url <- url[i]
  
  # Use tryCatch to capture errors and continue
  extracted_text <- tryCatch({
    # Normal execution: scrape the URL
    page <- read_html(current_url)
    page %>% html_nodes("p") %>% html_text()
  }, error = function(e) {
    # If something goes wrong, log the issue and return empty text
    message(paste("Skipping invalid URL:", current_url, "\nError:", e$message))
    character(0) # Returns an empty character vector; use "" if you want a single empty string
  })
  
  # Store the result for this URL
  results[[i]] <- extracted_text
}

# Optional: If you want to combine all results into a single vector
# combined_text <- unlist(results)

Key Improvements

  • Error Catching: tryCatch() lets us define exactly what happens when an error occurs—here, we log the problematic URL and return an empty value instead of crashing.
  • Proper Storage: Using a pre-allocated list (vector("list", length(url))) is way faster than dynamically growing a list, which is critical for 200k entries.
  • Transparency: The message() call lets you keep track of which URLs failed, so you can revisit them later if needed.

Bonus: Faster Parallel Processing (For 200k URLs)

A single loop will take forever with 200k URLs. For a speed boost, use the furrr package to run scrapes in parallel, paired with possibly() (a simpler alternative to tryCatch() for functional code):

library(rvest)
library(furrr)

# Set up parallel sessions (adjust based on your CPU cores)
plan(multisession, workers = 4)

# Define a "safe" version of your scraping function that returns empty on failure
safe_scrape_p <- possibly(function(link) {
  read_html(link) %>% html_nodes("p") %>% html_text()
}, otherwise = character(0))

# Run parallel scraping with a progress bar
results <- future_map(url, safe_scrape_p, .progress = TRUE)

This will cut down your runtime drastically—just make sure to add a small delay (like Sys.sleep(0.1)) inside the scraping function if you’re hitting a single domain, to avoid getting blocked by anti-scraping measures.

内容的提问来源于stack exchange,提问作者CatCaller

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 03:53:55