使用Selenium+Python在Chrome下载PDF:彻底禁用PDF Viewer问题求助
解决Chrome PDF Viewer框架残留及Selenium无法识别下载按钮问题
方案一:彻底禁用PDF Viewer,直接触发PDF下载
在现有Chrome配置基础上,补充更多偏好设置,强制Chrome直接下载PDF而非调用Viewer组件:
options = webdriver.ChromeOptions() options.add_experimental_option('prefs', { "download.default_directory": "file_path", # 修改下载路径 "download.prompt_for_download": False, "download.directory_upgrade": True, "plugins.always_open_pdf_externally": True, "plugins.plugins_disabled": ["Chrome PDF Viewer"], # 禁用PDF Viewer插件 "pdfjs.disabled": True, # 禁用内置PDF渲染器 "download.extensions_to_open": "" # 不自动打开任何下载文件 }) wd = webdriver.Chrome(options=options)
添加plugins.plugins_disabled和pdfjs.disabled参数后,Chrome会完全跳过PDF Viewer加载,直接触发下载,不会出现Viewer框架。
方案二:定位PDF Viewer内的下载按钮(若仍需保留Viewer)
如果必须保留Viewer框架,可通过Selenium定位Chrome内置的PDF Viewer元素——该元素处于Shadow DOM中,需通过JavaScript访问:
# 等待PDF Viewer加载完成 wd.implicitly_wait(10) # 执行JS获取Shadow DOM内的下载按钮并点击 download_button = wd.execute_script(""" return document.querySelector('embed').shadowRoot.querySelector('#download'); """) download_button.click()
注意:不同Chrome版本的PDF Viewer Shadow DOM结构可能略有差异,若上述选择器失效,可通过Chrome开发者工具的Elements面板查看实际选择器路径。
内容的提问来源于stack exchange,提问作者Austin Simpson
相关产品推荐
相关产品推荐

