亚马逊促销页爬虫脚本重复输出内容问题求助
Fixing Duplicate Content Issues in Your Amazon Deals Scraper
Hey there! Let's break down why your scraper is pulling duplicate content and fix it step by step.
What's Causing the Duplicates?
- Incorrect Selector Syntax: Your first
soup.select('div', {'class': ...})isn't working as intended. BeautifulSoup'sselect()method expects a CSS selector string as its first argument—passing the class as a second dictionary parameter doesn't filter the divs. This means you're looping through every div on the page, not just the deal containers you want. That's the core reason for duplicates. - Overly Nested Loops: Even if the selector worked, nesting three loops (div → button → a tag) increases the chance of reprocessing the same deal multiple times if elements are nested in unexpected ways.
- Scroll Execution Issue: Your
execute_scriptcall tries to do two things at once, which might not trigger the scroll properly. The scroll command should run separately before grabbing the page HTML.
Fixed Code
Here's the revised script with key changes explained:
# -*- coding: utf-8 -*- import time from bs4 import BeautifulSoup from selenium import webdriver # Initialize driver (note: for Selenium 4.x+, use Service instead of executable_path) driver = webdriver.Chrome(executable_path='C:\\Users\\Compu City\\Desktop\\chromedriver.exe') driver.get('https://www.amazon.com/international-sales-offers/b/?ie=UTF8&node=15529609011&ref_=nav_navm_intl_deal_btn') # Fix scroll execution: run scroll first, then wait for content to load time.sleep(3) # Adjust wait time based on your internet speed driver.execute_script("window.scrollBy(0, 2000);") time.sleep(2) # Give lazy-loaded content time to appear # Get page HTML after scrolling res = driver.execute_script("return document.documentElement.outerHTML") soup = BeautifulSoup(res, 'html5lib') # Use a set to automatically handle duplicates seen_titles = set() # Correct CSS selector to target only deal containers deal_containers = soup.select('div.a-section.a-spacing-none.tallCellView.gridColumn5.singleCell') for container in deal_containers: # Check if the container has the "Choose options" button button = container.select_one('button.a-button-text.a-text-center') if button and 'Choose options' in button.text.strip(): # Grab the deal title directly from the container title_tag = container.select_one('a.a-size-base.a-link-normal.dealTitleTwoLine.singleCellTitle.autoHeight') if title_tag: title = title_tag.text.strip() if title not in seen_titles: seen_titles.add(title) print(title) driver.close()
Key Improvements:
- Targeted Selectors: Used proper CSS class chaining to only fetch the deal containers you care about, not every div on the page.
- Simplified Logic: Replaced nested loops with
select_one()to grab individual elements directly from each container, cutting down redundant processing. - Built-in Duplicate Prevention: Added a
set()calledseen_titlesto store already printed titles, ensuring each deal is output only once. - Reliable Scroll: Separated the scroll command from HTML retrieval to make sure the page actually scrolls before content is captured.
- Error Resilience: Added checks to confirm elements exist before accessing their text, avoiding potential
AttributeErrors.
Bonus Tip for Newer Selenium Versions
If you're using Selenium 4.x+, the executable_path parameter is deprecated. Use the Service class instead:
from selenium.webdriver.chrome.service import Service service = Service('C:\\Users\\Compu City\\Desktop\\chromedriver.exe') driver = webdriver.Chrome(service=service)
内容的提问来源于stack exchange,提问作者colla bingo
相关产品推荐
相关产品推荐

