在R中使用rvest循环从data.frame批量抓取eBay商品数据
Hey Andrew, great question! Batch scraping multiple eBay India URLs in R is totally doable by wrapping your existing single-URL code into a reusable function and applying it across your URL list. Below’s a step-by-step solution with clean, efficient batch processing and error handling to avoid crashes if some URLs fail.
Step 1: Organize Your URLs into a Data Frame
First, gather all your target eBay URLs into a data frame. This makes it easy to manage and iterate over them:
# Create a data frame with your eBay URLs ebay_urls <- data.frame( url = c( "https://www.ebay.in/sch/i.html?_nkw=Mobile+Phones&_pgn=2&_skc=2&_skc=200&rt=nc", "https://www.ebay.in/sch/i.html?_nkw=Mobile+Phones&_pgn=2&_skc=10&_skc=1800&rt=nc", # Add all your remaining URLs here "https://www.ebay.in/sch/i.html?_nkw=Mobile+Phones&_pgn=3&_skc=5&_skc=1000&rt=nc" ), stringsAsFactors = FALSE )
Step 2: Build a Reusable Scraping Function
Turn your existing single-URL code into a function that handles errors gracefully. This ensures a single failed URL won’t break the entire batch:
library(rvest) library(stringr) library(purrr) # For streamlined batch processing # Define a function to scrape one eBay page scrape_ebay_page <- function(target_url) { tryCatch({ # Load the page HTML page <- read_html(target_url) # Extract data (adjust CSS selectors if eBay updates their page structure!) product_titles <- page %>% html_nodes(".s-item__title") %>% html_text() product_prices <- page %>% html_nodes(".s-item__price") %>% html_text() product_links <- page %>% html_nodes(".s-item__link") %>% html_attr("href") # Return a data frame of results for this page data.frame( Title = product_titles, Price = product_prices, Product_Link = product_links, Source_URL = target_url, # Track which URL the data came from stringsAsFactors = FALSE ) }, error = function(e) { # Print error details and return empty data frame for failed URLs message(paste("⚠️ Failed to scrape:", target_url, "\nError:", e$message)) return(data.frame(Title = character(), Price = character(), Product_Link = character(), Source_URL = character())) }) }
Note: Double-check the CSS selectors (like .s-item__title) by inspecting eBay’s page source—they might change over time.
Step 3: Batch Process All URLs
Use purrr::map_dfr to apply the function to every URL in your data frame and combine all results into one master data frame:
# Scrape all URLs and combine results all_ebay_data <- map_dfr(ebay_urls$url, scrape_ebay_page) # View the first few rows of your scraped data head(all_ebay_data)
Alternative: For Loop (If You Prefer)
If you’re more comfortable with for loops instead of purrr, here’s an equivalent approach:
# Initialize an empty data frame to store results all_ebay_data <- data.frame(Title = character(), Price = character(), Product_Link = character(), Source_URL = character()) # Loop through each URL for (url in ebay_urls$url) { page_data <- scrape_ebay_page(url) all_ebay_data <- rbind(all_ebay_data, page_data) }
content的提问来源于stack exchange,提问作者Andrew Seaman

