You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Selenium和Python根据前置文本或类名定位文本内容

解决方案:用Selenium根据前置文本或类名定位目标span

嘿,作为刚接触Web Scraping的新手,遇到这种动态类别的定位问题太正常了!别慌,我来给你分享两种靠谱的方法,完美适配你遇到的场景:


方法1:根据前置文本定位对应的span

这种方法适合前置文本(比如"First Category: ")比较稳定的情况,直接通过文本节点找到它后面的span元素。

核心思路

用XPath定位到包含目标前置文本的文本节点,然后通过following-sibling找到它后面紧邻的span兄弟元素。

代码示例(Python)

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
driver.get("你的目标网页URL")

# 定义要查找的前置文本
target_text = "First Category: "

try:
    # 等待元素加载完成,避免动态加载问题
    target_span = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located(
            (By.XPATH, f"//div[@class='div_class']/text()[contains(., '{target_text}')]/following-sibling::span[@class='span_class']")
        )
    )
    # 获取span的文本内容
    category_value = target_span.text
    print(f"{target_text}对应的值:{category_value}")
except Exception as e:
    # 处理类别缺失的情况
    print(f"未找到{target_text}对应的类别")

driver.quit()

解释XPath逻辑

  • //div[@class='div_class']:先定位到包含所有类别的父div
  • /text()[contains(., '{target_text}')]:找到div中包含目标前置文本的文本节点
  • /following-sibling::span[@class='span_class']:取这个文本节点后面的span兄弟元素

方法2:根据的类名定位对应的span

如果前置文本可能有变化(比如多语言、空格调整),但的类名是稳定的,这种方法更可靠。

核心思路

先定位到指定类名的元素,再通过following-sibling找到它后面的span元素(不管中间有没有文本节点,XPath会自动跳过文本节点找到span)。

代码示例(Python)

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
driver.get("你的目标网页URL")

# 定义要查找的<i>类名
target_i_class = "first_i_class"

try:
    target_span = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located(
            (By.XPATH, f"//div[@class='div_class']/i[@class='{target_i_class}']/following-sibling::span[@class='span_class']")
        )
    )
    category_value = target_span.text
    print(f"类名为{target_i_class}的<i>对应的span值:{category_value}")
except Exception as e:
    print(f"未找到类名为{target_i_class}的<i>对应的类别")

driver.quit()

额外提示

  1. 处理动态加载:一定要用WebDriverWait等待元素出现,避免因为页面还没加载完就查找元素导致的报错。
  2. 类别缺失的兼容:用try-except块捕获异常,这样即使某个类别缺失,程序也不会崩溃,可以返回默认值或者跳过该类别。
  3. XPath灵活性:如果网页结构有小变化,你可以调整XPath的逻辑,比如用contains()匹配部分类名(比如contains(@class, 'first_i')),应对类名有随机后缀的情况。

内容的提问来源于stack exchange,提问作者DRo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 11:18:09