如何使用Selenium获取包含指定文本元素的直接父节点类名
问题原因
- 原XPath表达式存在两个核心问题:
- 字符串拼接错误:关键词没有包裹引号,导致XPath语法非法
- 匹配逻辑问题:
contains(text(), 关键词)会匹配所有包含该文本的元素,父元素只要子孙节点有对应文本也会命中,find_element默认返回DOM树中最上层的匹配节点,因此拿到了顶层的html、外层container这类元素
- 原代码还存在变量名错误:定义的配置变量是
chrome_options,传入Chrome构造函数时写的是options,会直接报错。
解决方案
修改XPath表达式,增加过滤条件,仅匹配直接包含目标关键词的最内层元素,同时修正语法和变量错误,完整代码如下:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By chromeDriverPath = "./chromedriver" chrome_options = webdriver.ChromeOptions() # 修正变量名错误 driver = webdriver.Chrome(chromeDriverPath, options=chrome_options) driver.get("https://www.scrapethissite.com/pages/") # 要抓取的关键词列表 listOfKeywords = ['ajax', 'click'] for keyword in listOfKeywords: try: # 修正XPath:增加过滤条件,仅匹配最内层包含关键词的元素,同时给关键词加单引号 foundKeyword = driver.find_element(By.XPATH, f"//*[contains(text(), '{keyword}') and not(./*[contains(text(), '{keyword}')])]") print(foundKeyword.get_attribute("class")) except: pass driver.close()
运行结果
上述代码会输出你期望的结果:
page-title lead session-desc
内容的提问来源于stack exchange,提问作者Nick
相关产品推荐
相关产品推荐

