使用R抓取raworange服装站9页图片URL并下载的技术求助
Fixing Your Image Scraping & Download Script for Raw Orange's Clothing Collection
Let's break down why your original code wasn't working, then walk through a corrected version that handles all 9 pages and properly downloads images with their original filenames.
Issues with Your Original Code
- It only scrapes the first page of the collection—no logic to handle pagination for pages 2-9
html_node()only retrieves the first matching image instead of all images on the page; you needhtml_nodes()instead- The image
srcvalues are likely protocol-relative (like//cdn...) or partial URLs, so directly appending them to the collection page URL won't create a valid image link - No handling for potential anti-scraping measures (like missing a user-agent header, which can get you blocked)
Corrected Script
Here's a robust version that covers all 9 pages, fetches every product image, and downloads them with their original filenames:
library(rvest) library(purrr) # For cleaner iteration (optional but recommended) # 1. Generate URLs for all 9 pages base_url <- "https://www.raworange.com/collections/all-clothing" page_urls <- paste0(base_url, "?page=", 1:9) # 2. Add a user-agent to mimic a real browser (avoids being blocked) browser_headers <- add_headers( `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" ) # 3. Function to scrape all images from a single page scrape_single_page <- function(page_url) { # Fetch page content with the custom header page_content <- read_html(page_url, headers = browser_headers) # Extract all product image sources using a reliable selector img_sources <- page_content %>% html_nodes(xpath = '//div[contains(@class, "product-item")]//img') %>% html_attr("src") # Fix protocol-relative URLs (add https: to make them valid) img_sources <- ifelse(startsWith(img_sources, "//"), paste0("https:", img_sources), img_sources) return(img_sources) } # 4. Scrape images from all 9 pages all_image_urls <- map(page_urls, scrape_single_page) %>% unlist() # 5. Download each image with its original filename walk(all_image_urls, function(img_url) { # Extract the filename from the image URL image_filename <- basename(img_url) # Download the image (use mode = "wb" for binary files like images) download.file( url = img_url, destfile = image_filename, mode = "wb", quiet = FALSE ) # Add a 1-second delay to be respectful to the website's server Sys.sleep(1) })
Key Improvements Explained
- Pagination Handling: We generate URLs for all 9 pages and loop through each one to capture every product
- Fetch All Images:
html_nodes()(plural) grabs every image on the page instead of just the first match - URL Validation: Converts protocol-relative URLs to full HTTPS links so
download.file()can access them - Anti-Blocking Measures: A valid user-agent header helps the site recognize your request as a legitimate browser visit
- Binary Download:
mode = "wb"ensures images are saved correctly as binary files (not corrupted text) - Rate Limiting:
Sys.sleep(1)adds a small delay between downloads to avoid overwhelming the website's server
Important Notes
- Always check the website's
robots.txtfile before scraping to ensure you have permission to access and download their content - If the XPath selector stops working later, inspect the page's HTML to update the selector (class names on e-commerce sites can change over time)
内容的提问来源于stack exchange,提问作者ican. priyankanimavat
相关产品推荐
相关产品推荐

