You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium如何在不规范HTML中用XPath提取体育总监姓名邮箱

可精准定位体育总监信息的XPath表达式

提取姓名

//h6[contains(@class,'athletic-faculty-header') and text()='Athletic Director']/parent::div/following-sibling::strong[text()='Name:']/following-sibling::text()[1]

拿到结果后调用strip()方法去掉前后多余空白字符,就能得到干净的姓名文本。

提取邮箱地址

直接取邮箱文本

//h6[contains(@class,'athletic-faculty-header') and text()='Athletic Director']/parent::div/following-sibling::strong[text()='Email:']/following-sibling::a[1]/text()

取mailto协议链接(适用于需要处理链接的场景)

//h6[contains(@class,'athletic-faculty-header') and text()='Athletic Director']/parent::div/following-sibling::strong[text()='Email:']/following-sibling::a[1]/@href

表达式逻辑说明

这套表达式顺着页面层级做定位,完全避开了strong标签重复无法区分的问题:

  1. 先锁定「Athletic Director(体育总监)」对应的分类标题h6标签
  2. 找到这个标题所在的父级div块
  3. 只匹配这个div块后面紧跟着的姓名、邮箱字段内容
    不会匹配到其他职位的信息,精准度很高。

内容的提问来源于stack exchange,提问作者NoobDev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 06:09:03