如何修改Selenium类生成与Chrome检查工具一致的完整XPath
问题
尝试基于Selenium.WebElement生成完整XPath时,生成结果与谷歌浏览器检查工具的「复制完整XPath」不一致。例如在finance.yahoo.com历史价格数据页面中,Chrome生成的路径为:
/html/body/div[2]/main/section/section/section/article/section[1]/div[2]/div[1]/section/div/section[1]/div[2]
而自定义类生成的路径为:
/html/body/div[3]/main/section[1]/section[1]/section[2]/article/section[1]/div[14]/div[1]/section[1]/div[1]/section[1]/div[2]
目标是修改SeleniumXPath类,使其生成的XPath与Chrome工具结果匹配。原类代码如下:
Option Explicit Private pLabel As String Private pXPath As String Private pHasAnId As Boolean Private pTimeStamp As Date Public Property Let Label(aLabel As String) 'Set in factory pLabel = aLabel End Property Public Property Get Label() As String Label = pLabel End Property Public Property Let WebElement(anElem As Selenium.WebElement) 'Set in factory Call FindFullXPath(anElem) pTimeStamp = SNow End Property Public Property Get XPath() As String XPath = pXPath End Property Public Property Get hasAnId() As Boolean hasAnId = pHasAnId End Property Public Property Get IsStale() As Boolean If (SNow - pTimeStamp > 1) Then IsStale = True End If End Property Public Property Get Id() As String Id = XPath End Property Private Sub FindFullXPath(ByVal anElem As Selenium.WebElement) 'A helpful discussion can be found in ' https://www.youtube.com/watch?v=lr7CFZEI2YA Dim anId As String Dim finished As Boolean Dim aTag As String Dim ancestorElem As Selenium.WebElement Dim anXPath As String Dim anElemArr As Selenium.WebElements Dim aCount As Long Dim k2 As Long On Error Resume Next 'Attempt to retrieve the WebElement unique ID anId = anElem.Attribute("id") If (anId = "") Then 'A unique ID was not found so, from the given element ' traverse the html tree backwards to map the full path Do While (Not finished) aTag = anElem.tagName If (aTag = "html") Then finished = True Else Set ancestorElem = anElem.FindElementByXPath("..") If (anXPath <> "") Then anXPath = "/" & anXPath End If Set anElemArr = ancestorElem.FindElementsByTag(aTag) aCount = anElemArr.Count If (aCount > 1) Then For k2 = 1 To aCount If (anElemArr.item(k2).Equals(anElem)) Then Exit For End If Next k2 anXPath = aTag & "[" & k2 & "]" & anXPath Else anXPath = aTag & anXPath End If Set anElem = ancestorElem End If Loop pXPath = "/html/" & anXPath Else 'A unique id was found, use it instead of the full ' XPath pHasAnId = True pXPath = anId End If Call Err.Clear End Sub
解决方案
差异原因分析
- 子节点范围错误:原代码通过
FindElementsByTag获取父元素下所有层级的同标签后代元素,但Chrome的完整XPath仅统计父元素的直接子节点中同标签的数量,这是索引偏差的核心原因。 - 标签大小写不一致:Chrome生成的XPath标签全为小写,原代码直接使用
tagName返回的可能是大写(如HTML、BODY),导致路径格式不匹配。 - 元素匹配逻辑偏差:原代码遍历所有后代元素匹配目标节点,顺序可能与DOM文档的直接子节点顺序不一致。
修改后的代码
仅需替换原FindFullXPath方法,其余类成员保持不变:
Private Sub FindFullXPath(ByVal anElem As Selenium.WebElement) Dim anId As String Dim finished As Boolean Dim aTag As String Dim ancestorElem As Selenium.WebElement Dim anXPath As String Dim directChildren As Selenium.WebElements Dim childIndex As Long Dim currentChild As Selenium.WebElement On Error Resume Next ' 获取元素ID,优先使用标准ID XPath格式 anId = anElem.Attribute("id") If anId <> "" Then pHasAnId = True pXPath = "//*[@id='" & anId & "']" GoTo Cleanup End If ' 从目标元素向上遍历DOM树,构建完整XPath finished = False Do While Not finished aTag = LCase(anElem.tagName) ' 统一转小写,匹配Chrome格式 If aTag = "html" Then finished = True Else Set ancestorElem = anElem.FindElementByXPath("..") ' 获取父元素的所有直接子节点(而非所有后代) Set directChildren = ancestorElem.FindElementsByXPath("./child::*") ' 遍历直接子节点,找到当前元素的索引(从1开始) childIndex = 0 For Each currentChild In directChildren childIndex = childIndex + 1 If currentChild.Equals(anElem) Then Exit For End If Next currentChild ' 统计父元素下同标签的直接子节点数量,仅当数量>1时添加索引 Dim sameTagCount As Long sameTagCount = ancestorElem.FindElementsByXPath("./child::" & aTag).Count If sameTagCount > 1 Then anXPath = aTag & "[" & childIndex & "]/" & anXPath Else anXPath = aTag & "/" & anXPath End If Set anElem = ancestorElem End If Loop ' 拼接根节点,移除末尾多余的斜杠 pXPath = "/html/" & Left(anXPath, Len(anXPath) - 1) Cleanup: Call Err.Clear Set ancestorElem = Nothing Set directChildren = Nothing Set currentChild = Nothing End Sub
修改说明
- 直接子节点统计:使用
./child::*获取父元素的直接子节点,确保仅统计当前层级的同标签元素,和Chrome的逻辑一致。 - 标签小写转换:通过
LCase(anElem.tagName)将标签名统一转为小写,匹配Chrome的XPath格式。 - 索引逻辑优化:先找到当前元素在父节点直接子节点中的位置,再统计同标签直接子节点的数量,仅当数量大于1时添加索引,和Chrome的省略规则一致。
- ID处理优化:将原直接返回ID字符串的逻辑,修改为标准的ID XPath格式
//*[@id='xxx'],更符合XPath规范。
内容的提问来源于stack exchange,提问作者KC2
相关产品推荐
相关产品推荐

