You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python版Selenium:通过XPath获取不含子元素的节点文本

解决Selenium获取排除

标签的投资目标文本问题

嘿,我来帮你搞定这个Selenium获取文本的问题!根据你给出的页面结构,我们需要从那个fund-objective的div里提取排除

标题的投资目标文本,这里有几个靠谱的方案:

方案一:直接定位目标文本节点(最精准)

利用XPath直接筛选div下的非h3文本节点,一步到位拿到内容:

from selenium import webdriver
from selenium.webdriver.common.by import By

dr = webdriver.Chrome()
# 假设已打开目标页面,执行以下代码获取文本
objective_text = dr.find_element(
    By.XPATH, 
    "//div[@class='carousel-content column fund-objective']/text()[normalize-space()]"
).get_attribute('textContent').strip()

print(objective_text)

解释:

  • //div[@class='carousel-content column fund-objective']/text() 定位到div下所有直接子文本节点
  • [normalize-space()] 过滤掉空的文本节点(比如换行、多余空格)
  • get_attribute('textContent') 提取纯文本内容,最后用strip()清理首尾空格

方案二:先拿全部文本再移除标题(最直观)

如果文本结构简单,也可以先获取整个div的文本,再去掉h3的标题内容:

from selenium import webdriver
from selenium.webdriver.common.by import By

dr = webdriver.Chrome()
# 获取目标div元素
fund_div = dr.find_element(By.XPATH, "//div[@class='carousel-content column fund-objective']")
# 获取h3标题文本
h3_title = fund_div.find_element(By.TAG_NAME, 'h3').text
# 移除标题后得到目标文本
objective_text = fund_div.text.replace(h3_title, '').strip()

print(objective_text)

解释:

  • 先拿到div的全部文本,再用replace()去掉h3的标题部分
  • 最后用strip()处理掉替换后残留的多余空格,得到干净的内容

方案三:用JS执行XPath查询(适配复杂结构)

如果页面节点结构更复杂,也可以通过JS执行XPath查询,灵活筛选非h3节点的文本:

from selenium import webdriver

dr = webdriver.Chrome()
# 通过execute_script执行JS里的XPath逻辑
objective_text = dr.execute_script("""
return document.evaluate(
    "string(//div[@class='carousel-content column fund-objective']/node()[not(self::h3)])",
    document,
    null,
    XPathResult.STRING_TYPE,
    null
).stringValue;
""").strip()

print(objective_text)

解释:

  • node()[not(self::h3)] 选中div下所有不是h3的子节点
  • string() 将这些节点的文本拼接成字符串
  • 最后用strip()清理多余空格

内容的提问来源于stack exchange,提问作者P A N

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:16:57