You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取酒类网站:如何筛选同className下的目标商品元素?

解决Selenium爬取时筛选目标商品元素的问题

因为目标商品和底部新闻元素共用了catalog_product_item_cont类名,你可以通过以下几种方式精准筛选出商品元素:

1. 限定查找范围到商品父容器

先定位商品列表所在的专属父容器,再在这个容器内查找目标元素,直接排除新闻板块的内容。需要先通过浏览器开发者工具(F12)确认商品列表的父容器类名或标识。

示例代码:

# 假设商品列表的父容器类名为catalog_product_list(需根据实际页面调整)
product_container = Driver.find_element(By.CLASS_NAME, "catalog_product_list")
# 仅在商品父容器内查找目标元素
items = product_container.find_elements(By.CLASS_NAME, "catalog_product_item_cont")

print(f"{len(items)} items found")
for item in items:
    print(item.text)

2. 通过元素特征筛选

商品元素必然包含价格、容量这类特有信息,而新闻元素没有。可以遍历所有匹配元素,检查是否存在商品特有的子元素或文本特征。

示例代码:

from selenium.common.exceptions import NoSuchElementException

items = Driver.find_elements(By.CLASS_NAME,"catalog_product_item_cont")
valid_products = []

for item in items:
    try:
        # 检查元素是否包含价格子元素(假设价格元素类名为price)
        item.find_element(By.CLASS_NAME, "price")
        # 额外验证:检查文本是否包含容量单位(如ml、L)
        if "ml" in item.text or "L" in item.text:
            valid_products.append(item)
    except NoSuchElementException:
        # 无价格元素,判定为新闻内容,跳过
        continue

print(f"{len(valid_products)} valid items found")
for product in valid_products:
    print(product.text)

3. 使用精准的选择器

利用CSS选择器或XPath,直接指定目标元素的层级关系或附加属性,区分商品和新闻元素。

CSS选择器示例

# 假设商品元素在.product_catalog容器下,新闻元素不在该容器内
items = Driver.find_elements(By.CSS_SELECTOR, ".product_catalog .catalog_product_item_cont")

XPath示例

# 定位有catalog_product_item_cont类名,且父级包含product_catalog类的元素
items = Driver.find_elements(By.XPATH, "//div[@class='catalog_product_item_cont' and ancestor::div[@class='product_catalog']]")

以上方法的核心是利用页面DOM结构的差异(父容器、子元素、层级关系)来区分两类元素,你需要根据目标网站的实际结构选择最合适的方案。

内容的提问来源于stack exchange,提问作者Лев Поваляев

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 19:40:38