You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R加速网页爬取?需2小时内完成23000页数据抓取

Optimizing R Web Scraping for 23k Pages in 2 Hours

Hey there! Let's tackle this web scraping speed issue together—23k pages in 2 hours is totally doable with some key optimizations, especially since your core scraping logic is already solid. Let's break down the fixes step by step:

Core Bottlenecks in Your Current Code

First, let's identify what's slowing your script down:

  • Repeated rbind() in loops: Every time you append to tabellafinale, R copies the entire data frame. For 23k iterations, this becomes an enormous waste of time.
  • Single-threaded execution: You're only scraping one page at a time, leaving most of your machine's processing power unused.
  • Redundant code: You’re repeating the same gsub clean-up steps three times, and creating unnecessary data.table objects per iteration.
  • Basic error handling: The try() block works, but we can make it more flexible without breaking the whole scrape.

Key Optimization Steps & Tools

1. Switch to Parallel Processing

The biggest speed gain will come from scraping multiple pages at once. We’ll use furrr (a beginner-friendly wrapper for parallel workflows) since it plays nicely with tidyverse tools you might already be familiar with.

First, install and load the necessary packages:

install.packages(c("furrr", "rvest", "data.table", "purrr"))
library(furrr)
library(rvest)
library(data.table)
library(purrr)

Set up parallel sessions (adjust workers based on your CPU cores—4-6 is a safe starting point):

plan(multisession, workers = 6)

2. Wrap Scraping Logic in a Reusable Function

Let’s condense your scraping and cleaning steps into a single function. This makes it easy to apply in parallel and keeps your code clean.

scrape_single_page <- function(code) {
  # Build the URL
  url <- paste0("https://sample/", code)
  
  # Use `possibly()` to handle errors gracefully without stopping the whole scrape
  scrape_result <- possibly(function() {
    # Fetch and parse the page
    page <- read_html(url)
    
    # Extract raw text from selectors
    carr_raw <- html_text(html_nodes(page, ".span3"))
    prez_raw <- html_text(html_nodes(page, ".carbFormat"))
    viag_raw <- html_text(html_nodes(page, ".span5"))
    
    # Create a helper function to clean text (avoids repeating gsub calls)
    clean_text <- function(text) {
      text |>
        gsub("\n", "", x = _) |>
        gsub("\r", "", x = _) |>
        gsub("\t", "", x = _)
    }
    
    # Clean all text fields
    carr_clean <- clean_text(carr_raw)
    prez_clean <- clean_text(prez_raw)
    viag_clean <- clean_text(viag_raw)
    
    # Extract the relevant subset of carr (from index 7 onwards)
    carb_final <- carr_clean[7:length(carr_clean)]
    
    # Return a data.table with results for this page
    data.table(
      gest = code,
      cr = carb_final,
      pr = prez_clean,
      vig = viag_clean
    )
  }, otherwise = data.table(gest = code, cr = NA, pr = NA, vig = NA))()
  
  return(scrape_result)
}

3. Run Parallel Scraping & Combine Results

Instead of looping, use future_map_dfr to apply the function to all codes in parallel. It will automatically combine results into a single data frame (which we can convert to data.table):

# Get your list of codes
imp <- Codes$id

# Run the scrape with a progress bar (so you can track progress!)
tabellafinale <- future_map_dfr(imp, scrape_single_page, .progress = TRUE)

# Convert to data.table (optional, but matches your original workflow)
setDT(tabellafinale)

Extra Tips for Speed & Reliability

  • Optimize HTTP Requests: Use httr2 instead of read_html directly to add timeouts, custom user agents (to avoid being blocked), and retries for failed requests. Here’s a quick tweak to the scrape function:
    library(httr2)
    
    scrape_single_page <- function(code) {
      url <- paste0("https://sample/", code)
      
      scrape_result <- possibly(function() {
        resp <- request(url) |>
          req_user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") |>
          req_timeout(10) |>  # Time out after 10 seconds if no response
          req_perform()
        
        page <- resp |> resp_body_html()
        # Rest of the scraping logic stays the same...
      }, otherwise = data.table(gest = code, cr = NA, pr = NA, vig = NA))()
      
      return(scrape_result)
    }
    
  • Avoid Overloading the Server: Don’t set too many parallel workers—this can get your IP blocked. Start with 4-6 and adjust if you see errors or slowdowns.
  • Check for Hidden APIs: Many websites load data via JSON APIs instead of rendering it directly in HTML. Use your browser’s DevTools (Network tab) to look for API requests—scraping JSON is way faster than parsing HTML!

Final Notes

With these changes, you should easily hit your 2-hour target. Parallel processing alone will speed up your workflow by 4-6x, and the other tweaks will reduce overhead and improve reliability. Test with a small subset of codes first to work out any kinks!

内容的提问来源于stack exchange,提问作者G_scrape7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 20:07:35