求助:如何用Selenium抓取#shadow-root (open)下的href标签
抓取Shadow DOM中的href标签问题解决
问题背景
作为编程新手,需要抓取网站https://iltacon2022.expofp.com/中#shadow-root (open)下隐藏的href标签。尝试使用Selenium代码时遇到NoSuchElementException错误,期望提取出类似?access-corp、?accruent-inc这类格式的href内容。
原代码:
from bs4 import BeautifulSoup as bs from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time url = "https://iltacon2022.expofp.com/" options = webdriver.ChromeOptions() options.add_argument("start-maximized") options.add_experimental_option("detach", True) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()),options=options) driver.get(url) time.sleep(6) root1 = driver.find_element(By.XPATH,"/html/body/div[1]/div").shadow_root root2 = driver.find_element(By.XPATH,"/html/body/div[1]/div//div") print(root2) driver.quit()
错误信息:
selenium.common.exceptions.NoSuchElementException: Message: no such element: Unable to locate element: {"method":"xpath","selector":"/html/body/div[1]/div//div"}
错误原因
- Shadow DOM隔离限制:Shadow DOM内的元素无法通过
driver.find_element()直接定位,必须通过对应的shadow_root对象来查找内部元素。原代码获取root1后仍用driver查找root2,自然找不到Shadow DOM里的元素。 - 等待方式不可靠:
time.sleep(6)是固定时长等待,无法确保元素实际加载完成,容易因页面加载延迟导致报错。
修正后的代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = "https://iltacon2022.expofp.com/" options = webdriver.ChromeOptions() options.add_argument("start-maximized") driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) driver.get(url) # 显式等待包含shadow-root的宿主元素加载完成 shadow_host = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "body > div:first-child > div")) ) # 获取Shadow DOM的根节点 shadow_root = shadow_host.shadow_root # 在Shadow DOM内查找所有带href属性的a标签 links = WebDriverWait(shadow_root, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "a[href]")) ) # 提取href中?开头的部分并打印 for link in links: href = link.get_attribute("href") if "?" in href: print(href.split("?")[-1]) driver.quit()
代码说明
- 显式等待:用
WebDriverWait替代固定等待,确保元素加载完成后再执行操作,避免因页面加载慢导致的元素未找到问题。 - Shadow DOM访问:通过
shadow_host.shadow_root获取Shadow DOM的根节点,后续所有内部元素的查找都基于这个根节点进行。 - href提取:遍历找到的a标签,提取href中
?之后的部分,完全匹配期望输出格式。
内容的提问来源于stack exchange,提问作者BQuist
相关产品推荐
相关产品推荐

