You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

动态XPath爬取商品数据遇阻,寻求非XPath爬取替代方案

问题描述

尝试从Cdiscount网站爬取商品名称、价格及其他信息时遇到以下问题:

  • 商品名称的XPath仅单个标签变化,较易获取,但商品价格的XPath存在多处标签差异(比如1号商品价格XPath是//*[@id="lpBloc"]/li[1]/div[2]/div[3]/div[1]/div/div[2]/span[1],2号是//*[@id="lpBloc"]/li[2]/div[2]/div[2]/div[1]/div/div[2]/span[1])
  • 直接通过类名和ID爬取价格时返回空字符串
  • 希望找到无需依赖绝对XPath的替代爬取方法

原代码如下:

driver= webdriver.Chrome('E:/chromedriver/chromedriver.exe')
product_name=[]
product_price=[]
product_rating=[]
product_url=[]
driver.get('https://www.cdiscount.com/bricolage/climatisation/traitement-de-l-air/ioniseur/l-166130303.html#_his_')
for i in range(1,55):
    try :
        productname=driver.find_element('xpath','//*[@id="lpBloc"]/li['+str(i)+']/a/div[2]/div/span').text
        product_name.append(productname)
    except:
        print("none")
print(product_name)
替代爬取方法

1. 基于商品容器的相对定位

先定位到每个商品的父容器(#lpBloc下的每个li),再在容器内部查找目标元素,摆脱绝对XPath的依赖:

from selenium import webdriver

driver = webdriver.Chrome('E:/chromedriver/chromedriver.exe')
product_name = []
product_price = []
product_rating = []
product_url = []

driver.get('https://www.cdiscount.com/bricolage/climatisation/traitement-de-l-air/ioniseur/l-166130303.html#_his_')

# 获取所有商品的父容器
product_items = driver.find_elements('css selector', '#lpBloc li')

for item in product_items:
    # 爬取商品名称
    try:
        name = item.find_element('css selector', 'a div[class^="productName"] span').text
        product_name.append(name)
    except:
        product_name.append(None)
    
    # 爬取商品价格:优先匹配带价格特征的元素
    try:
        # 匹配包含price关键词的类或带有data-price属性的元素
        price = item.find_element('css selector', '[class*="price"], [data-price]').text
        product_price.append(price)
    except:
        # 备选方案:筛选包含欧元符号的文本元素
        price_elems = item.find_elements('css selector', 'span')
        for elem in price_elems:
            if '€' in elem.text:
                product_price.append(elem.text)
                break
        else:
            product_price.append(None)

print("商品名称:", product_name)
print("商品价格:", product_price)

2. 利用文本特征的相对XPath

如果价格元素必然包含€符号,可在商品容器内用相对路径查找,不用写绝对层级:

# 在单个商品容器内查找价格
price = item.find_element('xpath', './/span[contains(text(), "€")]').text

这里的.表示从当前商品容器节点开始查找,避免了标签层级变化的影响。

3. 直接解析页面内嵌JSON数据

多数电商网站会把商品数据嵌入页面的<script>标签中,直接解析这些JSON可以完全跳过DOM定位:

import json
from selenium import webdriver

driver = webdriver.Chrome('E:/chromedriver/chromedriver.exe')
driver.get('https://www.cdiscount.com/bricolage/climatisation/traitement-de-l-air/ioniseur/l-166130303.html#_his_')

# 遍历页面所有script标签,查找包含商品数据的内容
scripts = driver.find_elements('css selector', 'script')
for script in scripts:
    content = script.get_attribute('innerHTML')
    if 'productList' in content or 'products' in content:
        # 截取并解析JSON部分(需根据实际内容调整截取逻辑)
        start_idx = content.find('{')
        end_idx = content.rfind('}') + 1
        product_data = json.loads(content[start_idx:end_idx])
        # 提取商品信息
        for product in product_data.get('products', []):
            product_name.append(product.get('name'))
            product_price.append(product.get('price'))
        break

print("商品名称:", product_name)
print("商品价格:", product_price)

注意事项

  • 爬取前需遵守网站robots.txt规则,避免触发反爬机制
  • 可添加显式等待(WebDriverWait)替代固定等待,确保元素加载完成
  • 若遇到动态生成的类名,优先使用属性包含选择器([class*="xxx"])

内容的提问来源于stack exchange,提问作者Lord of the Strings

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 10:50:41