You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

亚马逊促销页爬虫脚本重复输出内容问题求助

Fixing Duplicate Content Issues in Your Amazon Deals Scraper

Hey there! Let's break down why your scraper is pulling duplicate content and fix it step by step.

What's Causing the Duplicates?

  1. Incorrect Selector Syntax: Your first soup.select('div', {'class': ...}) isn't working as intended. BeautifulSoup's select() method expects a CSS selector string as its first argument—passing the class as a second dictionary parameter doesn't filter the divs. This means you're looping through every div on the page, not just the deal containers you want. That's the core reason for duplicates.
  2. Overly Nested Loops: Even if the selector worked, nesting three loops (div → button → a tag) increases the chance of reprocessing the same deal multiple times if elements are nested in unexpected ways.
  3. Scroll Execution Issue: Your execute_script call tries to do two things at once, which might not trigger the scroll properly. The scroll command should run separately before grabbing the page HTML.

Fixed Code

Here's the revised script with key changes explained:

# -*- coding: utf-8 -*-
import time
from bs4 import BeautifulSoup
from selenium import webdriver

# Initialize driver (note: for Selenium 4.x+, use Service instead of executable_path)
driver = webdriver.Chrome(executable_path='C:\\Users\\Compu City\\Desktop\\chromedriver.exe')
driver.get('https://www.amazon.com/international-sales-offers/b/?ie=UTF8&node=15529609011&ref_=nav_navm_intl_deal_btn')

# Fix scroll execution: run scroll first, then wait for content to load
time.sleep(3)  # Adjust wait time based on your internet speed
driver.execute_script("window.scrollBy(0, 2000);")
time.sleep(2)  # Give lazy-loaded content time to appear

# Get page HTML after scrolling
res = driver.execute_script("return document.documentElement.outerHTML")
soup = BeautifulSoup(res, 'html5lib')

# Use a set to automatically handle duplicates
seen_titles = set()

# Correct CSS selector to target only deal containers
deal_containers = soup.select('div.a-section.a-spacing-none.tallCellView.gridColumn5.singleCell')

for container in deal_containers:
    # Check if the container has the "Choose options" button
    button = container.select_one('button.a-button-text.a-text-center')
    if button and 'Choose options' in button.text.strip():
        # Grab the deal title directly from the container
        title_tag = container.select_one('a.a-size-base.a-link-normal.dealTitleTwoLine.singleCellTitle.autoHeight')
        if title_tag:
            title = title_tag.text.strip()
            if title not in seen_titles:
                seen_titles.add(title)
                print(title)

driver.close()

Key Improvements:

  • Targeted Selectors: Used proper CSS class chaining to only fetch the deal containers you care about, not every div on the page.
  • Simplified Logic: Replaced nested loops with select_one() to grab individual elements directly from each container, cutting down redundant processing.
  • Built-in Duplicate Prevention: Added a set() called seen_titles to store already printed titles, ensuring each deal is output only once.
  • Reliable Scroll: Separated the scroll command from HTML retrieval to make sure the page actually scrolls before content is captured.
  • Error Resilience: Added checks to confirm elements exist before accessing their text, avoiding potential AttributeErrors.

Bonus Tip for Newer Selenium Versions

If you're using Selenium 4.x+, the executable_path parameter is deprecated. Use the Service class instead:

from selenium.webdriver.chrome.service import Service

service = Service('C:\\Users\\Compu City\\Desktop\\chromedriver.exe')
driver = webdriver.Chrome(service=service)

内容的提问来源于stack exchange,提问作者colla bingo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 23:27:32