Selenium如何获取含子标签元素文本并过滤<i>标签内容
解决Selenium获取标签文本并排除内部标签文本的问题
你遇到的报错是因为findElement()要求返回元素节点,但你的XPath表达式//div[@class='cart-content-btn']//a/i/following-sibling::node()匹配到的是文本节点,不符合方法的返回要求。下面给出两种可靠的解决方法:
方法一:通过JavaScript直接提取目标文本
这种方法直接操作DOM节点,精准过滤掉<i>元素的文本,只保留<a>下的纯文本节点内容:
Java代码示例
// 先定位到目标<a>元素 WebElement aElement = driver.findElement(By.xpath("//div[@class='cart-content-btn']//a")); // 执行JavaScript提取排除<i>后的文本 String targetText = (String) ((JavascriptExecutor) driver).executeScript(""" const aEl = arguments[0]; let result = ''; // 遍历<a>的所有子节点 for (const node of aEl.childNodes) { // 只保留非空的文本节点 if (node.nodeType === Node.TEXT_NODE && node.textContent.trim()) { result += node.textContent.trim(); } } return result; """, aElement);
Python代码示例
from selenium.webdriver.common.by import By # 定位<a>元素 a_element = driver.find_element(By.XPATH, "//div[@class='cart-content-btn']//a") # 执行JS脚本 target_text = driver.execute_script(""" const aEl = arguments[0]; let result = ''; for (const node of aEl.childNodes) { if (node.nodeType === Node.TEXT_NODE && node.textContent.trim()) { result += node.textContent.trim(); } } return result; """, a_element)
方法二:通过文本替换间接获取(适合简单场景)
如果<a>内只有一个<i>元素,且<i>的文本和<a>的其他文本无重叠,可以先获取<a>的完整文本,再减去<i>的文本:
Java代码示例
WebElement aElement = driver.findElement(By.xpath("//div[@class='cart-content-btn']//a")); WebElement iElement = aElement.findElement(By.tagName("i")); String fullText = aElement.getText().trim(); String iText = iElement.getText().trim(); // 替换掉<i>的文本并去除多余空格 String targetText = fullText.replace(iText, "").trim();
Python代码示例
a_element = driver.find_element(By.XPATH, "//div[@class='cart-content-btn']//a") i_element = a_element.find_element(By.TAG_NAME, "i") full_text = a_element.text.strip() i_text = i_element.text.strip() target_text = full_text.replace(i_text, "").strip()
注意:方法二的局限性在于如果
<i>的文本和<a>的其他文本有重复内容,替换会出错,因此优先推荐方法一。
内容的提问来源于stack exchange,提问作者Ip Man
相关产品推荐
相关产品推荐

