You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用BeautifulSoup做Web爬虫时重复输出商品如何去重

爬虫重复产品输出去重方案

实现思路

利用Python集合的元素唯一性特性,记录已经输出过的产品标识,仅首次出现的产品会被打印,后续重复出现的同款产品自动过滤。

修正后完整代码

from bs4 import BeautifulSoup
import requests

page = requests.get("https://www.lamazuna.com/en/")
soup = BeautifulSoup(page.content, "html.parser")

all_product_item_lists = soup.find_all(class_="col-sm-12 mega-col")
# 初始化集合存储已打印的产品标识
seen_products = set()

for product_link in all_product_item_lists:
    for link in product_link.find_all("a", href=True):
        find_product_url = link.get('href')

        next_page = requests.get(find_product_url)
        next_soup = BeautifulSoup(next_page.content, "html.parser")

        product_name_none = next_soup.find(class_="h3 product-title")
        product_price_none = next_soup.find(class_="price")

        if product_name_none is not None:
            product_name = product_name_none.get_text().strip()
            # 仅当产品未被记录时执行打印逻辑
            if product_name not in seen_products:
                if product_price_none is not None:
                    product_price = product_price_none.get_text().strip()
                # 可按需调整打印内容,这里保留原需求仅打印产品名
                print(product_name)
                seen_products.add(product_name)

核心修改说明

  • 循环外初始化空集合seen_products,用于存储已经打印过的产品名称
  • 每次获取到有效产品名后,先判断是否存在于集合中,不存在才执行打印操作,同时将该产品名加入集合
  • 优化了原代码的变量作用域逻辑,避免产品信息为空时出现变量未定义的报错
  • 用strip()替代replace("\n",""),可以同时清除首尾的空格、换行、制表符等冗余字符,处理更彻底

如果担心存在同名不同款的产品,可以将去重判断依据换成产品的唯一URL,只需要把集合存储的内容从产品名改为find_product_url即可,去重准确性更高。

内容的提问来源于stack exchange,提问作者Henrik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 12:54:03