Python使用BeautifulSoup做Web爬虫时重复输出商品如何去重
爬虫重复产品输出去重方案
实现思路
利用Python集合的元素唯一性特性,记录已经输出过的产品标识,仅首次出现的产品会被打印,后续重复出现的同款产品自动过滤。
修正后完整代码
from bs4 import BeautifulSoup import requests page = requests.get("https://www.lamazuna.com/en/") soup = BeautifulSoup(page.content, "html.parser") all_product_item_lists = soup.find_all(class_="col-sm-12 mega-col") # 初始化集合存储已打印的产品标识 seen_products = set() for product_link in all_product_item_lists: for link in product_link.find_all("a", href=True): find_product_url = link.get('href') next_page = requests.get(find_product_url) next_soup = BeautifulSoup(next_page.content, "html.parser") product_name_none = next_soup.find(class_="h3 product-title") product_price_none = next_soup.find(class_="price") if product_name_none is not None: product_name = product_name_none.get_text().strip() # 仅当产品未被记录时执行打印逻辑 if product_name not in seen_products: if product_price_none is not None: product_price = product_price_none.get_text().strip() # 可按需调整打印内容,这里保留原需求仅打印产品名 print(product_name) seen_products.add(product_name)
核心修改说明
- 循环外初始化空集合
seen_products,用于存储已经打印过的产品名称 - 每次获取到有效产品名后,先判断是否存在于集合中,不存在才执行打印操作,同时将该产品名加入集合
- 优化了原代码的变量作用域逻辑,避免产品信息为空时出现变量未定义的报错
- 用
strip()替代replace("\n",""),可以同时清除首尾的空格、换行、制表符等冗余字符,处理更彻底
如果担心存在同名不同款的产品,可以将去重判断依据换成产品的唯一URL,只需要把集合存储的内容从产品名改为find_product_url即可,去重准确性更高。
内容的提问来源于stack exchange,提问作者Henrik
相关产品推荐
相关产品推荐

