使用Selenium爬取ZALORA遇重复数据及无法跳转下页问题求助
Fixing Duplicate Data & Pagination Issues in Your ZALORA Selenium Scraper
Hey there! Let’s tackle those two frustrating issues you’re facing with your ZALORA scraper—duplicate data being pulled every time, and getting stuck on the first page instead of moving to the next. I’ll walk you through the likely causes and give you a revised code example that fixes both problems.
Why You’re Getting Duplicate Data
Chances are, your current code is either:
- Reusing old element references from the first page instead of fetching fresh elements for each new page
- Grabbing data before the page has fully loaded, leading to stale or repeated content
- Using global variables that don’t reset between page loops, causing old data to persist
Why Pagination Isn’t Working
Common culprits here include:
- Incorrectly targeting the "Next Page" button (wrong selector, or trying to click it before it’s clickable)
- Not waiting for the new page to load after clicking, so your scraper keeps pulling the same page’s data
- No check for when there are no more pages to scrape, leading to infinite loops or errors
Revised Code with Fixes
Here’s an updated version of your script that addresses both problems. I’ve added comments to explain key changes:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # Initialize driver and wait object (explicit waits are more reliable than implicit) driver = webdriver.Chrome() wait = WebDriverWait(driver, 30) # Fixed URL (replaced & with actual & for proper formatting) base_url = 'https://www.zalora.com.hk/men/clothing/shirt/?gender=men&dir=desc&sort=popularity&category_id=31&page={}&enable_visual_sort=1' current_page = 1 try: while True: print(f"Scraping page {current_page}...") # Load the current page driver.get(base_url.format(current_page)) # Wait for product items to fully load (adjust selector to match ZALORA's actual HTML) product_items = wait.until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.product-item")) ) # Loop through each product on the page to avoid duplicates for item in product_items: try: # Extract brand and title by searching within the product item (not globally) brand_name = item.find_element(By.CSS_SELECTOR, "div.brand-name").text.strip() product_title = item.find_element(By.CSS_SELECTOR, "div.product-title").text.strip() print(f"Brand: {brand_name} | Title: {product_title}") # Add code here to save data to a list, CSV, etc. except Exception as e: print(f"Failed to extract item: {str(e)}") continue # Check for next page and navigate try: # Wait for next page button to be clickable (adjust selector to match ZALORA's button) next_button = wait.until( EC.element_to_be_clickable((By.CSS_SELECTOR, "a.pagination__next")) ) # Stop if next button is disabled (some sites add a 'disabled' class to inactive buttons) if "disabled" in next_button.get_attribute("class"): print("No more pages available. Exiting...") break # Click next page and wait for the new page to load next_button.click() wait.until(EC.url_contains(f"page={current_page + 1}")) current_page += 1 time.sleep(2) # Small buffer to avoid overwhelming the site except Exception as e: print(f"Could not navigate to next page: {str(e)}") break finally: # Always close the driver when done driver.quit()
Key Fixes Explained
For Duplicate Data:
- Fresh Element Fetching: Each page load uses
wait.untilto pull new product elements, so you never reuse old references. - Local Extraction: Brand and title are extracted from within each individual product item (
item.find_element), ensuring you only get data from the current page’s items. - No Global Variables: Data is stored temporarily per item instead of global variables that hold old values.
For Pagination:
- Explicit Wait for Clickable Button: Ensures the next page button is fully loaded and clickable before trying to interact with it.
- Disabled Button Check: Stops the loop when there are no more pages to scrape.
- URL Wait: Confirms the page has actually navigated to the next page before proceeding to scrape.
Important Notes
- Adjust Selectors: Use your browser’s developer tools (F12) to verify the actual CSS selectors for product items, brand names, titles, and the next page button—ZALORA’s HTML might change over time.
- Respect Site Rules: Add reasonable delays between requests and avoid scraping too aggressively to avoid being blocked by ZALORA’s anti-bot measures.
- Handle Edge Cases: Add more error handling if needed (e.g., timeouts, missing elements) to make your scraper more robust.
内容的提问来源于stack exchange,提问作者TedLLH
相关产品推荐
相关产品推荐

