使用LXML/Requests-HTML定位同父元素下纯文本后的span元素
网页抓取定位解决方案
基于lxml的XPath方案
利用XPath的文本节点定位和兄弟节点选择,精准锁定目标<span>:
- 核心思路:先找到包含固定纯文本的文本节点,再选取它紧随其后的
<span>元素 - 示例代码:
from lxml import etree # 假设已获取网页内容并解析为HTML树 html_content = "你的网页HTML内容" tree = etree.HTML(html_content) # 替换CONSISTENT TEXT:为实际固定文本,normalize-space用于处理空白/换行 target_span = tree.xpath("//td/text()[normalize-space()='CONSISTENT TEXT:']/following-sibling::span[1]") if target_span: desired_info = target_span[0].text.strip() print(desired_info)
也可以反向匹配:找<span>节点,其前一个兄弟文本节点是目标固定文本:
target_span = tree.xpath("//td/span[preceding-sibling::text()[normalize-space()='CONSISTENT TEXT:']]")
基于requests-html的实现
requests-html原生支持XPath,直接复用上述逻辑即可:
from requests_html import HTMLSession session = HTMLSession() response = session.get("目标网页URL") # 同样使用XPath定位 target_span = response.html.xpath("//td/text()[normalize-space()='CONSISTENT TEXT:']/following-sibling::span[1]") if target_span: desired_info = target_span[0].text.strip() print(desired_info)
关键说明
normalize-space():处理文本中的空格、换行、制表符,避免因格式问题导致匹配失败following-sibling::span[1]:确保只选取固定文本后第一个<span>,避免匹配到后续无关的<span>
内容的提问来源于stack exchange,提问作者amota
相关产品推荐
相关产品推荐

