You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Selenium爬取亚马逊子分类下全部页面商品?

Solution to Scrape All Pages of Amazon Subcategory

Here's how you can modify your script to loop through all 400+ pages and extract product links from each one:

Modified Code

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import random
import time

# Proxy setup (keep your existing proxy)
PROXY = "88.157.149.250:8080"
chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument(f'--proxy-server={PROXY}')

# Updated XPath for product links (using modern Selenium syntax)
LINKS_XPATH = '//*[contains(@id,"result")]/div/div[3]/div[1]/a'
# XPath for "Next Page" button (adjust if Amazon's UI changes)
NEXT_BUTTON_XPATH = '//a[@class="s-pagination-item s-pagination-next s-pagination-button s-pagination-separator"]'

# Initialize browser
browser = webdriver.Chrome(options=chrome_options)
browser.get('https://www.amazon.com/s/ref=lp_11444071011_nr_p_8_1/132-3636705-4291947?rh=n%3A3375251%2Cn%3A%213375301%2Cn%3A10971181011%2Cn%3A11444071011%2Cp_8%3A2229059011')

try:
    has_next_page = True
    page_number = 1
    
    while has_next_page:
        print(f"Scraping page {page_number}...")
        
        # Wait for product links to load on current page
        WebDriverWait(browser, 10).until(
            EC.presence_of_all_elements_located((By.XPATH, LINKS_XPATH))
        )
        
        # Extract all product links
        links = browser.find_elements(By.XPATH, LINKS_XPATH)
        for link in links:
            href = link.get_attribute('href')
            print(href)
        
        # Try to navigate to next page
        try:
            # Wait for next button to be clickable
            next_button = WebDriverWait(browser, 5).until(
                EC.element_to_be_clickable((By.XPATH, NEXT_BUTTON_XPATH))
            )
            # Add random delay to mimic human behavior
            time.sleep(random.uniform(1, 3))
            next_button.click()
            page_number += 1
        except:
            # No more next page or button is not clickable
            has_next_page = False
            print("No more pages to scrape.")
            
finally:
    # Close the browser when done
    browser.quit()

Key Changes Explained

  • Loop Structure: Added a while loop that runs until there's no more "Next Page" button, ensuring we cover all available pages.
  • Modern Selenium Syntax: Replaced deprecated find_elements_by_xpath with find_elements(By.XPATH, ...)—this is required for Selenium 4+ compatibility.
  • Explicit Waits: Used WebDriverWait to ensure elements are fully loaded before interacting with them, preventing errors from slow page loads or dynamic content.
  • Random Delays: Added time.sleep(random.uniform(1,3)) to mimic human browsing speed, which helps avoid triggering Amazon's anti-scraping measures.
  • Error Handling: Wrapped the next button interaction in a try-except block to gracefully exit the loop when there are no more pages to scrape.
  • Cleanup: Added a finally block to ensure the browser is closed even if an error occurs during scraping.

Important Notes

  • Amazon's Anti-Scraping: Amazon actively blocks scrapers. Even with a proxy, you might encounter captchas or temporary blocks. Consider using rotating proxies and adding more human-like behavior (e.g., random scrolls, mouse movements) if you run into issues.
  • UI Changes: Amazon occasionally updates its website structure. If the "Next Page" button XPath stops working, inspect the button element to get the updated class or attribute.
  • Rate Limiting: Be mindful of how many requests you send in a short time. Too many rapid requests will get you blocked.

内容的提问来源于stack exchange,提问作者ryy77

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:41:50