如何用Python Selenium点击ShadowRoot中PDF阅读器的下载按钮
Selenium 操作Chrome内置PDF下载按钮失败问题处理
问题场景
- 使用Python Selenium访问链接
https://cissearch.kcc.gov.tw/System/Bulletin/View.aspx?BulletinSN=239928&pages=9957#pdfStart,目标是自动点击PDF阅读器的下载按钮 - 试过Stack Overflow的下载配置方案,仍需手动点击页面上的
open按钮才能触发下载 - 参考LambdaTest的Shadow DOM定位方法,无法成功定位元素
- 执行以下JS代码时,浏览器控制台能正常运行,但Python中抛出错误:
driver.execute_script("document.querySelector('pdf-viewer').shadowRoot.querySelector('viewer-toolbar').shadowRoot.querySelector('viewer-download-controls').shadowRoot.querySelector('cr-action-menu').querySelector('button')")
错误信息:
JavascriptException: Message: javascript error: Cannot read properties of null (reading 'shadowRoot') (Session info: chrome=109.0.5414.87)
核心原因
Python中执行JS时,页面的PDF嵌套组件(尤其是多层Shadow DOM结构)可能还未完全加载完成,导致querySelector返回null,进而无法读取shadowRoot。而浏览器控制台能正常运行,是因为手动操作时页面已经完全加载完毕。
解决方案
方案1:添加显式等待+空值判断,确保元素加载
先等待最外层的pdf-viewer元素出现,再逐层定位Shadow DOM内的元素,同时增加空值判断避免报错:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 等待pdf-viewer元素加载完成,最长等待20秒 WebDriverWait(driver, 20).until( EC.presence_of_element_located((By.TAG_NAME, "pdf-viewer")) ) # 执行带空值判断的JS代码定位下载按钮 download_btn = driver.execute_script(""" // 逐层定位,每一步都做空值校验 const pdfViewer = document.querySelector('pdf-viewer'); if (!pdfViewer?.shadowRoot) return null; const toolbar = pdfViewer.shadowRoot.querySelector('viewer-toolbar'); if (!toolbar?.shadowRoot) return null; const downloadControls = toolbar.shadowRoot.querySelector('viewer-download-controls'); if (!downloadControls?.shadowRoot) return null; const actionMenu = downloadControls.shadowRoot.querySelector('cr-action-menu'); return actionMenu?.querySelector('button') || null; """) if download_btn: download_btn.click() else: print("未找到PDF下载按钮")
方案2:配置Chrome自动下载PDF,跳过阅读器操作
直接让Chrome自动下载PDF,无需打开内置阅读器,省去点击下载按钮的步骤:
from selenium import webdriver from selenium.webdriver.chrome.options import Options chrome_options = Options() # 配置Chrome自动下载PDF chrome_options.add_experimental_option('prefs', { "download.default_directory": "/your/local/download/path", # 替换成你的本地下载路径 "download.prompt_for_download": False, "download.directory_upgrade": True, "plugins.always_open_pdf_externally": True, # 直接下载,不打开内置PDF阅读器 "plugins.plugins_list": [{"enabled": False, "name": "Chrome PDF Viewer"}] }) driver = webdriver.Chrome(options=chrome_options) driver.get("https://cissearch.kcc.gov.tw/System/Bulletin/View.aspx?BulletinSN=239928&pages=9957#pdfStart")
注:如果页面是通过跳转加载PDF,可能需要先定位页面上的PDF链接元素,直接点击链接触发下载。
方案3:直接提取PDF链接下载(最稳定)
绕过Selenium操作浏览器的环节,从页面源码中提取PDF的真实链接,用requests直接下载:
import requests import re # 获取页面源码 page_source = driver.page_source # 从源码中匹配PDF链接(根据页面实际结构调整正则) pdf_url_pattern = r'https?://[^\s"]+\.pdf' pdf_urls = re.findall(pdf_url_pattern, page_source) if pdf_urls: pdf_url = pdf_urls[0] # 发送请求下载PDF response = requests.get(pdf_url) with open("kcc_bulletin.pdf", "wb") as f: f.write(response.content) print("PDF已成功下载") else: print("未找到PDF下载链接")
内容的提问来源于stack exchange,提问作者Hackore
相关产品推荐
相关产品推荐

