动态XPath爬取商品数据遇阻,寻求非XPath爬取替代方案
问题描述
尝试从Cdiscount网站爬取商品名称、价格及其他信息时遇到以下问题:
- 商品名称的XPath仅单个标签变化,较易获取,但商品价格的XPath存在多处标签差异(比如1号商品价格XPath是
//*[@id="lpBloc"]/li[1]/div[2]/div[3]/div[1]/div/div[2]/span[1],2号是//*[@id="lpBloc"]/li[2]/div[2]/div[2]/div[1]/div/div[2]/span[1]) - 直接通过类名和ID爬取价格时返回空字符串
- 希望找到无需依赖绝对XPath的替代爬取方法
原代码如下:
driver= webdriver.Chrome('E:/chromedriver/chromedriver.exe') product_name=[] product_price=[] product_rating=[] product_url=[] driver.get('https://www.cdiscount.com/bricolage/climatisation/traitement-de-l-air/ioniseur/l-166130303.html#_his_') for i in range(1,55): try : productname=driver.find_element('xpath','//*[@id="lpBloc"]/li['+str(i)+']/a/div[2]/div/span').text product_name.append(productname) except: print("none") print(product_name)
替代爬取方法
1. 基于商品容器的相对定位
先定位到每个商品的父容器(#lpBloc下的每个li),再在容器内部查找目标元素,摆脱绝对XPath的依赖:
from selenium import webdriver driver = webdriver.Chrome('E:/chromedriver/chromedriver.exe') product_name = [] product_price = [] product_rating = [] product_url = [] driver.get('https://www.cdiscount.com/bricolage/climatisation/traitement-de-l-air/ioniseur/l-166130303.html#_his_') # 获取所有商品的父容器 product_items = driver.find_elements('css selector', '#lpBloc li') for item in product_items: # 爬取商品名称 try: name = item.find_element('css selector', 'a div[class^="productName"] span').text product_name.append(name) except: product_name.append(None) # 爬取商品价格:优先匹配带价格特征的元素 try: # 匹配包含price关键词的类或带有data-price属性的元素 price = item.find_element('css selector', '[class*="price"], [data-price]').text product_price.append(price) except: # 备选方案:筛选包含欧元符号的文本元素 price_elems = item.find_elements('css selector', 'span') for elem in price_elems: if '€' in elem.text: product_price.append(elem.text) break else: product_price.append(None) print("商品名称:", product_name) print("商品价格:", product_price)
2. 利用文本特征的相对XPath
如果价格元素必然包含€符号,可在商品容器内用相对路径查找,不用写绝对层级:
# 在单个商品容器内查找价格 price = item.find_element('xpath', './/span[contains(text(), "€")]').text
这里的.表示从当前商品容器节点开始查找,避免了标签层级变化的影响。
3. 直接解析页面内嵌JSON数据
多数电商网站会把商品数据嵌入页面的<script>标签中,直接解析这些JSON可以完全跳过DOM定位:
import json from selenium import webdriver driver = webdriver.Chrome('E:/chromedriver/chromedriver.exe') driver.get('https://www.cdiscount.com/bricolage/climatisation/traitement-de-l-air/ioniseur/l-166130303.html#_his_') # 遍历页面所有script标签,查找包含商品数据的内容 scripts = driver.find_elements('css selector', 'script') for script in scripts: content = script.get_attribute('innerHTML') if 'productList' in content or 'products' in content: # 截取并解析JSON部分(需根据实际内容调整截取逻辑) start_idx = content.find('{') end_idx = content.rfind('}') + 1 product_data = json.loads(content[start_idx:end_idx]) # 提取商品信息 for product in product_data.get('products', []): product_name.append(product.get('name')) product_price.append(product.get('price')) break print("商品名称:", product_name) print("商品价格:", product_price)
注意事项
- 爬取前需遵守网站
robots.txt规则,避免触发反爬机制 - 可添加显式等待(
WebDriverWait)替代固定等待,确保元素加载完成 - 若遇到动态生成的类名,优先使用属性包含选择器(
[class*="xxx"])
内容的提问来源于stack exchange,提问作者Lord of the Strings
相关产品推荐
相关产品推荐

