使用Selenium迭代时如何访问每个元素的非唯一子类?
实现方法
你可以通过当前商品节点的相对XPath查找实现精准匹配,不会和其他商品的同class图片元素混淆,具体操作如下:
方法1:先定位商品节点再查找子元素(更灵活,方便同时拿商品文本和图片)
你拿到每一个序号对应的商品节点后,直接在该节点的上下文范围内查找图片元素即可,注意XPath开头加.代表从当前节点开始检索,不要从全局文档检索:
示例代码(Selenium环境):
from selenium import webdriver from selenium.webdriver.common.by import By import requests driver = webdriver.Chrome() driver.get("你的目标页面URL") # 替换为实际的商品总数量 goods_total = 20 for x in range(1, goods_total+1): # 定位当前序号的商品节点 current_goods_xpath = f'//*[@id="app"]/div[1]/main/div[2]/section/div/section/div[4]/div[1]/div[{x}]' current_goods = driver.find_element(By.XPATH, current_goods_xpath) # 获取商品名称+价格文本,后续可自行拆分字段 goods_info = current_goods.text.strip().split("\n") # 根据实际文本换行规则调整 goods_name, goods_price = goods_info[0], goods_info[1] # 查找当前商品下的图片元素 img_ele = current_goods.find_element(By.XPATH, './/*[contains(@class, "shop-card__image-block")]') # 提取图片链接,懒加载页面可替换为data-src、data-original等属性名 img_url = img_ele.get_attribute("src") # 下载图片示例 img_content = requests.get(img_url).content with open(f"{goods_name}_{goods_price}.jpg", "wb") as f: f.write(img_content)
方法2:直接拼接全局XPath获取图片
也可以直接把图片路径拼到商品XPath后面,直接获取对应序号商品的图片元素:
//*[@id="app"]/div[1]/main/div[2]/section/div/section/div[4]/div[1]/div[X]//*[contains(@class, "shop-card__image-block")]
把X替换为1~n的对应序号即可直接定位到目标图片元素。
注意事项
- 如果页面是懒加载渲染,需要先滚动到对应商品的位置,再获取图片属性,否则可能拿到占位图的无效链接
- 如果你确认
shop-card__image-block是元素的唯一class,可以把contains(@class, "shop-card__image-block")替换为@class="shop-card__image-block"
内容的提问来源于stack exchange,提问作者Georgije Tanasic
相关产品推荐
相关产品推荐

