使用Python Selenium点击Chrome下载按钮无法下载PDF压缩包的问题
问题解决方案
以下是针对你遇到的下载未触发问题的具体解决步骤:
1. 修正元素定位逻辑
你当前点击的是按钮内部的<span>标签,这可能无法触发按钮的点击事件。直接定位到<button>标签本身:
button = driver.find_element(By.XPATH, '//*[@id="react-root"]/div/div/div/main/section[1]/div/div/div/form/button')
2. 完善Chrome下载偏好配置
添加更多针对下载安全限制、PDF插件的设置,避免拦截:
options.add_experimental_option('prefs', { "download.default_directory": r"你的完整下载路径", "download.prompt_for_download": False, "download.directory_upgrade": True, "plugins.always_open_pdf_externally": True, "safebrowsing.enabled": False, "safebrowsing.disable_download_protection": True, "plugins.plugins_list": [{"enabled": False, "name": "Chrome PDF Viewer"}], "download.extensions_to_open": "" })
注意:确保下载路径存在且有读写权限,Windows路径使用原始字符串(r"路径")或双反斜杠。
3. 使用显式等待替代直接定位
避免因元素未加载完成导致点击失效,用WebDriverWait等待元素可点击:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 最长等待10秒,直到按钮可点击 button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="react-root"]/div/div/div/main/section[1]/div/div/div/form/button')) ) button.click()
4. 检查网站权限限制
ScienceDirect的部分内容需要登录订阅账号才能下载。先手动打开页面点击下载按钮,确认是否需要登录。如果需要,在Selenium代码中添加登录步骤(如填写账号密码)。
5. 替换固定sleep为下载完成等待
time.sleep(5)可能不足以等待下载完成,新增函数检查下载目录中的文件状态:
import os def wait_for_download(download_dir, timeout=30): end_time = time.time() + timeout while time.time() < end_time: for filename in os.listdir(download_dir): # 排除Chrome临时下载文件(.crdownload后缀) if not filename.endswith('.crdownload'): return True time.sleep(1) return False # 点击按钮后调用 wait_for_download(r"你的完整下载路径")
修改后的完整代码
from selenium import webdriver import time import os from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from webdriver_manager.chrome import ChromeDriverManager # 配置下载目录 download_dir = r"C:\你的完整下载路径" os.makedirs(download_dir, exist_ok=True) options = Options() options.add_experimental_option('prefs', { "download.default_directory": download_dir, "download.prompt_for_download": False, "download.directory_upgrade": True, "plugins.always_open_pdf_externally": True, "safebrowsing.enabled": False, "safebrowsing.disable_download_protection": True, "plugins.plugins_list": [{"enabled": False, "name": "Chrome PDF Viewer"}], "download.extensions_to_open": "" }) service = Service(ChromeDriverManager().install()) driver = webdriver.Chrome(service=service, options=options) driver.get('https://www.sciencedirect.com/journal/international-journal-of-greenhouse-gas-control/vol/127/suppl/C') # 如需登录,在此添加登录逻辑 # driver.find_element(By.ID, 'username').send_keys('你的账号') # driver.find_element(By.ID, 'password').send_keys('你的密码') # driver.find_element(By.ID, 'submit-btn').click() # 等待按钮可点击并触发下载 button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="react-root"]/div/div/div/main/section[1]/div/div/div/form/button')) ) button.click() # 等待下载完成 def wait_for_download(download_dir, timeout=30): end_time = time.time() + timeout while time.time() < end_time: files = os.listdir(download_dir) if files: for f in files: if not f.endswith('.crdownload'): print(f"下载完成:{f}") return True time.sleep(1) print("下载超时") return False wait_for_download(download_dir) driver.quit()
内容的提问来源于stack exchange,提问作者user22139764
相关产品推荐
相关产品推荐

