You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest的html_nodes()无法抓取亚马逊商品详情的问题

Fixing Your Amazon India Smartphone Scraper with rvest

Hey there! I see you're stuck trying to pull product details from Amazon India's search results using rvest—let's break down the issues and fix this step by step.

1. Your CSS Selector Is Incorrect

The main roadblock here is how you're targeting the product name span. When selecting an element with multiple classes, you need to chain the class names with dots (.), not spaces. Spaces in CSS selectors mean "find a child element with this class"—your original code was looking for a span with class a-size-medium, then a child element with class a-color-base, then another child with a-text-normal, which doesn't match the page's actual structure.

Wrong Selector:

"span.a-size-medium a-color-base a-text-normal"

Correct Selector:

"span.a-size-medium.a-color-base.a-text-normal"

2. Amazon's Anti-Scraping Measures

Amazon actively blocks default web scraper requests (including rvest's default read_html() call). You need to mimic a real browser by adding a User-Agent header to your request to avoid being flagged as a bot.

Fixed Full Code

Here's the revised code that addresses both issues, plus extracts all three requested details (name, price, stock status) into a tidy data frame:

library(rvest)
library(httr)
library(dplyr)
library(purrr)

# Set a realistic User-Agent (mimics a Chrome browser on Windows)
ua <- "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"

# Fetch the page with proper browser-like headers
url <- "https://www.amazon.in/s?k=Smartphone&rh=n%3A1389401031&ref=nb_sb_noss"
response <- GET(url, user_agent(ua))
page <- read_html(response)

# Extract all product parent containers first (to keep related details paired)
product_containers <- page %>% html_nodes("div.s-result-item")

# Scrape details for each individual product
product_data <- map_df(product_containers, function(container) {
  tibble(
    name = container %>% 
      html_node("span.a-size-medium.a-color-base.a-text-normal") %>% 
      html_text(trim = TRUE) %>% 
      replace_na("Product name unavailable"),
    price = container %>% 
      html_node("span.a-price-whole") %>% 
      html_text(trim = TRUE) %>% 
      replace_na("Price unavailable"),
    stock_status = container %>% 
      html_node("span.a-size-small.a-color-success") %>% 
      html_text(trim = TRUE) %>% 
      replace_na("Out of stock or status not listed")
  )
})

# View the final scraped data
print(product_data)

Key Notes:

  • User-Agent Flexibility: If the provided UA string stops working, you can grab your actual browser's UA via about:version (Chrome) or about:support (Firefox) and update it.
  • Container Targeting: Using the parent div.s-result-item ensures we pair each product's name, price, and stock status correctly—no risk of mismatched data from pulling all names then all prices separately.
  • Missing Value Handling: replace_na() keeps your data frame intact even if some products lack a detail (e.g., out-of-stock items might not display a price).

内容的提问来源于stack exchange,提问作者sp29

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 20:37:46