如何用Python+Selenium爬取亚马逊子分类下全部页面商品?
Solution to Scrape All Pages of Amazon Subcategory
Here's how you can modify your script to loop through all 400+ pages and extract product links from each one:
Modified Code
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import random import time # Proxy setup (keep your existing proxy) PROXY = "88.157.149.250:8080" chrome_options = webdriver.ChromeOptions() chrome_options.add_argument(f'--proxy-server={PROXY}') # Updated XPath for product links (using modern Selenium syntax) LINKS_XPATH = '//*[contains(@id,"result")]/div/div[3]/div[1]/a' # XPath for "Next Page" button (adjust if Amazon's UI changes) NEXT_BUTTON_XPATH = '//a[@class="s-pagination-item s-pagination-next s-pagination-button s-pagination-separator"]' # Initialize browser browser = webdriver.Chrome(options=chrome_options) browser.get('https://www.amazon.com/s/ref=lp_11444071011_nr_p_8_1/132-3636705-4291947?rh=n%3A3375251%2Cn%3A%213375301%2Cn%3A10971181011%2Cn%3A11444071011%2Cp_8%3A2229059011') try: has_next_page = True page_number = 1 while has_next_page: print(f"Scraping page {page_number}...") # Wait for product links to load on current page WebDriverWait(browser, 10).until( EC.presence_of_all_elements_located((By.XPATH, LINKS_XPATH)) ) # Extract all product links links = browser.find_elements(By.XPATH, LINKS_XPATH) for link in links: href = link.get_attribute('href') print(href) # Try to navigate to next page try: # Wait for next button to be clickable next_button = WebDriverWait(browser, 5).until( EC.element_to_be_clickable((By.XPATH, NEXT_BUTTON_XPATH)) ) # Add random delay to mimic human behavior time.sleep(random.uniform(1, 3)) next_button.click() page_number += 1 except: # No more next page or button is not clickable has_next_page = False print("No more pages to scrape.") finally: # Close the browser when done browser.quit()
Key Changes Explained
- Loop Structure: Added a
whileloop that runs until there's no more "Next Page" button, ensuring we cover all available pages. - Modern Selenium Syntax: Replaced deprecated
find_elements_by_xpathwithfind_elements(By.XPATH, ...)—this is required for Selenium 4+ compatibility. - Explicit Waits: Used
WebDriverWaitto ensure elements are fully loaded before interacting with them, preventing errors from slow page loads or dynamic content. - Random Delays: Added
time.sleep(random.uniform(1,3))to mimic human browsing speed, which helps avoid triggering Amazon's anti-scraping measures. - Error Handling: Wrapped the next button interaction in a
try-exceptblock to gracefully exit the loop when there are no more pages to scrape. - Cleanup: Added a
finallyblock to ensure the browser is closed even if an error occurs during scraping.
Important Notes
- Amazon's Anti-Scraping: Amazon actively blocks scrapers. Even with a proxy, you might encounter captchas or temporary blocks. Consider using rotating proxies and adding more human-like behavior (e.g., random scrolls, mouse movements) if you run into issues.
- UI Changes: Amazon occasionally updates its website structure. If the "Next Page" button XPath stops working, inspect the button element to get the updated class or attribute.
- Rate Limiting: Be mindful of how many requests you send in a short time. Too many rapid requests will get you blocked.
内容的提问来源于stack exchange,提问作者ryy77
相关产品推荐
相关产品推荐

