You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R抓取raworange服装站9页图片URL并下载的技术求助

Fixing Your Image Scraping & Download Script for Raw Orange's Clothing Collection

Let's break down why your original code wasn't working, then walk through a corrected version that handles all 9 pages and properly downloads images with their original filenames.

Issues with Your Original Code

  • It only scrapes the first page of the collection—no logic to handle pagination for pages 2-9
  • html_node() only retrieves the first matching image instead of all images on the page; you need html_nodes() instead
  • The image src values are likely protocol-relative (like //cdn...) or partial URLs, so directly appending them to the collection page URL won't create a valid image link
  • No handling for potential anti-scraping measures (like missing a user-agent header, which can get you blocked)

Corrected Script

Here's a robust version that covers all 9 pages, fetches every product image, and downloads them with their original filenames:

library(rvest)
library(purrr) # For cleaner iteration (optional but recommended)

# 1. Generate URLs for all 9 pages
base_url <- "https://www.raworange.com/collections/all-clothing"
page_urls <- paste0(base_url, "?page=", 1:9)

# 2. Add a user-agent to mimic a real browser (avoids being blocked)
browser_headers <- add_headers(
  `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
)

# 3. Function to scrape all images from a single page
scrape_single_page <- function(page_url) {
  # Fetch page content with the custom header
  page_content <- read_html(page_url, headers = browser_headers)
  
  # Extract all product image sources using a reliable selector
  img_sources <- page_content %>%
    html_nodes(xpath = '//div[contains(@class, "product-item")]//img') %>%
    html_attr("src")
  
  # Fix protocol-relative URLs (add https: to make them valid)
  img_sources <- ifelse(startsWith(img_sources, "//"), paste0("https:", img_sources), img_sources)
  
  return(img_sources)
}

# 4. Scrape images from all 9 pages
all_image_urls <- map(page_urls, scrape_single_page) %>% unlist()

# 5. Download each image with its original filename
walk(all_image_urls, function(img_url) {
  # Extract the filename from the image URL
  image_filename <- basename(img_url)
  
  # Download the image (use mode = "wb" for binary files like images)
  download.file(
    url = img_url,
    destfile = image_filename,
    mode = "wb",
    quiet = FALSE
  )
  
  # Add a 1-second delay to be respectful to the website's server
  Sys.sleep(1)
})

Key Improvements Explained

  • Pagination Handling: We generate URLs for all 9 pages and loop through each one to capture every product
  • Fetch All Images: html_nodes() (plural) grabs every image on the page instead of just the first match
  • URL Validation: Converts protocol-relative URLs to full HTTPS links so download.file() can access them
  • Anti-Blocking Measures: A valid user-agent header helps the site recognize your request as a legitimate browser visit
  • Binary Download: mode = "wb" ensures images are saved correctly as binary files (not corrupted text)
  • Rate Limiting: Sys.sleep(1) adds a small delay between downloads to avoid overwhelming the website's server

Important Notes

  • Always check the website's robots.txt file before scraping to ensure you have permission to access and download their content
  • If the XPath selector stops working later, inspect the page's HTML to update the selector (class names on e-commerce sites can change over time)

内容的提问来源于stack exchange,提问作者ican. priyankanimavat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:21:53