You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Selenium点击ShadowRoot中PDF阅读器的下载按钮

Selenium 操作Chrome内置PDF下载按钮失败问题处理

问题场景

  • 使用Python Selenium访问链接 https://cissearch.kcc.gov.tw/System/Bulletin/View.aspx?BulletinSN=239928&pages=9957#pdfStart,目标是自动点击PDF阅读器的下载按钮
  • 试过Stack Overflow的下载配置方案,仍需手动点击页面上的open按钮才能触发下载
  • 参考LambdaTest的Shadow DOM定位方法,无法成功定位元素
  • 执行以下JS代码时,浏览器控制台能正常运行,但Python中抛出错误:
driver.execute_script("document.querySelector('pdf-viewer').shadowRoot.querySelector('viewer-toolbar').shadowRoot.querySelector('viewer-download-controls').shadowRoot.querySelector('cr-action-menu').querySelector('button')")

错误信息:

JavascriptException: Message: javascript error: Cannot read properties of null (reading 'shadowRoot')
  (Session info: chrome=109.0.5414.87)

核心原因

Python中执行JS时,页面的PDF嵌套组件(尤其是多层Shadow DOM结构)可能还未完全加载完成,导致querySelector返回null,进而无法读取shadowRoot。而浏览器控制台能正常运行,是因为手动操作时页面已经完全加载完毕。

解决方案

方案1:添加显式等待+空值判断,确保元素加载

先等待最外层的pdf-viewer元素出现,再逐层定位Shadow DOM内的元素,同时增加空值判断避免报错:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 等待pdf-viewer元素加载完成,最长等待20秒
WebDriverWait(driver, 20).until(
    EC.presence_of_element_located((By.TAG_NAME, "pdf-viewer"))
)

# 执行带空值判断的JS代码定位下载按钮
download_btn = driver.execute_script("""
    // 逐层定位,每一步都做空值校验
    const pdfViewer = document.querySelector('pdf-viewer');
    if (!pdfViewer?.shadowRoot) return null;
    
    const toolbar = pdfViewer.shadowRoot.querySelector('viewer-toolbar');
    if (!toolbar?.shadowRoot) return null;
    
    const downloadControls = toolbar.shadowRoot.querySelector('viewer-download-controls');
    if (!downloadControls?.shadowRoot) return null;
    
    const actionMenu = downloadControls.shadowRoot.querySelector('cr-action-menu');
    return actionMenu?.querySelector('button') || null;
""")

if download_btn:
    download_btn.click()
else:
    print("未找到PDF下载按钮")

方案2:配置Chrome自动下载PDF,跳过阅读器操作

直接让Chrome自动下载PDF,无需打开内置阅读器,省去点击下载按钮的步骤:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

chrome_options = Options()
# 配置Chrome自动下载PDF
chrome_options.add_experimental_option('prefs', {
    "download.default_directory": "/your/local/download/path",  # 替换成你的本地下载路径
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
    "plugins.always_open_pdf_externally": True,  # 直接下载,不打开内置PDF阅读器
    "plugins.plugins_list": [{"enabled": False, "name": "Chrome PDF Viewer"}]
})

driver = webdriver.Chrome(options=chrome_options)
driver.get("https://cissearch.kcc.gov.tw/System/Bulletin/View.aspx?BulletinSN=239928&pages=9957#pdfStart")

注:如果页面是通过跳转加载PDF,可能需要先定位页面上的PDF链接元素,直接点击链接触发下载。

方案3:直接提取PDF链接下载(最稳定)

绕过Selenium操作浏览器的环节,从页面源码中提取PDF的真实链接,用requests直接下载:

import requests
import re

# 获取页面源码
page_source = driver.page_source

# 从源码中匹配PDF链接(根据页面实际结构调整正则)
pdf_url_pattern = r'https?://[^\s"]+\.pdf'
pdf_urls = re.findall(pdf_url_pattern, page_source)

if pdf_urls:
    pdf_url = pdf_urls[0]
    # 发送请求下载PDF
    response = requests.get(pdf_url)
    with open("kcc_bulletin.pdf", "wb") as f:
        f.write(response.content)
    print("PDF已成功下载")
else:
    print("未找到PDF下载链接")

内容的提问来源于stack exchange,提问作者Hackore

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 19:10:19