使用Selenium提取多个同名class下首个h3标签返回空数组如何解决
问题解决方案
已知错误点
- XPath的节点索引从1开始计算,你写的
[0]不会匹配到任何元素,直接导致返回空数组 elementor-widget-container是h3标签外层容器的类名,不是h3自身的类属性,匹配规则本身无法命中目标元素- 你直接将WebElement对象存入字典,即使匹配到元素也无法得到预期的名称文本
- 当前代码未做页面加载等待,可能元素还未渲染完成就执行了查找逻辑,你之前的WebDriverWait未生效大概率是定位规则写错导致的
- 你使用的
find_elements_by_xpath、executable_path参数属于Selenium 3.x的旧API,在4.0+版本中已被弃用,存在兼容性风险
修正后可用代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = 'https://www.riversidemedgroup.com/riverside-urgent-care/' data = [] # 替换为你的chromedriver路径 driver_path = '/Library/Frameworks/Python.framework/Versions/3.9/bin/chromedriver' driver = webdriver.Chrome(service=Service(driver_path)) driver.get(url) # 等待所有外层容器加载完成 containers = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, "elementor-widget-container")) ) for container in containers: try: # 取每个容器下的第一个h3标签 first_h3 = container.find_element(By.XPATH, "./h3[1]") data.append({ "Center Name": first_h3.text.strip() }) # 跳过没有h3的容器 except: continue print(data) driver.close()
额外优化建议
- 如果页面存在大量无关的
elementor-widget-container容器,可以先缩小定位范围,找到医疗中心列表对应的父级容器后再遍历查找,减少无效遍历 - 可以通过类名直接定位h3标签,你可以打开浏览器控制台元素面板,查看目标h3的专属类名后替换定位规则,效率更高
内容的提问来源于stack exchange,提问作者VRapport
相关产品推荐
相关产品推荐

