AWS EC2中Python Selenium+Firefox下载PDF文件失败问题
AWS EC2上Selenium Firefox无法保存PDF到指定目录的问题排查与解决
问题背景
开发Python项目时,通过Selenium+Firefox自动从网站下载PDF并上传至存储桶:
- 开发环境基于Docker容器,搭配Selenium 3.141.0/3.8.0/4.18.1版本与Firefox驱动,PDF可正常保存至
/tmp目录 - 部署到AWS EC2实例后,应用与网站交互全程无报错,指定的
unique_dir目录已生成,机器剩余21GB可用空间,但目录始终为空,PDF文件无处可寻
可能原因
- Firefox无头模式限制:EC2环境下,无头模式的Firefox可能因沙箱机制或
/dev/shm资源限制,导致下载路径配置不生效 - MIME类型不匹配:网站返回的PDF文件Content-Type可能不是
application/pdf,Firefox未触发自动保存逻辑(无头模式下无法弹出保存对话框) - 版本兼容性问题:EC2上的Firefox或Geckodriver版本与开发环境不一致,导致下载偏好配置失效
- 等待逻辑缺陷:原逻辑仅检查
.pdf后缀文件,但下载过程中会生成临时文件(如.part),可能因等待时间不足或未识别临时文件导致超时
解决方案
1. 调整Firefox无头模式配置
添加适配EC2环境的参数,补充下载相关偏好:
--no-sandbox:禁用沙箱,避免权限限制--disable-dev-shm-usage:绕过/dev/shm的大小限制- 强制指定下载目录、禁用内置PDF查看器,避免在线打开PDF
2. 扩展MIME类型覆盖范围
将常见的PDF相关MIME类型都加入自动保存列表,覆盖application/octet-stream、application/x-pdf等可能的类型
3. 优化下载等待逻辑
增加对临时下载文件(.part)的检测,延长等待时间,确保文件完成下载后再退出循环
4. 排查日志与环境一致性
- 查看
/tmp/geckodriver.log中的下载相关日志,定位权限或路径错误 - 确保EC2上的Firefox、Geckodriver版本与开发环境Docker内完全一致
修改后的代码示例
import os import stat import uuid import time import logging from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.firefox.options import Options from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC login_url = 'https://my-web.com/login/' dashboard_url = f'https://my-web.com/pdf-button-download-view' unique_dir = os.path.join("/tmp/pdfs", str(uuid.uuid4())) os.makedirs(unique_dir, exist_ok=True) # 确保目录权限足够 os.chmod(unique_dir, stat.S_IRWXU | stat.S_IRWXG | stat.S_IRWXO) options = Options() options.headless = True # 适配EC2无头环境的关键参数 options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") options.add_argument("--disable-gpu") log_path = "/tmp/geckodriver.log" # 配置Firefox下载偏好 profile = webdriver.FirefoxProfile() profile.set_preference("browser.download.folderList", 2) profile.set_preference("browser.download.manager.showWhenStarting", False) profile.set_preference("browser.download.dir", unique_dir) # 覆盖所有可能的PDF MIME类型 profile.set_preference("browser.helperApps.neverAsk.saveToDisk", "application/pdf,application/octet-stream,application/x-pdf") # 强制使用指定下载目录 profile.set_preference("browser.download.useDownloadDir", True) # 禁用内置PDF查看器,强制下载 profile.set_preference("pdfjs.disabled", True) # 关闭下载管理器相关弹窗(无头模式无法处理) profile.set_preference("browser.download.manager.useWindow", False) profile.set_preference("browser.download.manager.focusWhenStarting", False) profile.set_preference("browser.download.manager.closeWhenDone", True) driver = webdriver.Firefox(options=options, firefox_profile=profile, service_log_path=log_path) try: driver.get(login_url) time.sleep(10) # 登录流程 driver.find_element(By.ID, "username").send_keys("admin") driver.find_element(By.ID, "password").send_keys("admin") driver.find_element(By.CSS_SELECTOR, "form").submit() time.sleep(10) # 进入仪表盘 driver.get(dashboard_url) time.sleep(10) logging.info(f"Dashboard URL loaded: {driver.current_url}") # 触发下拉菜单 dropdown_trigger = driver.find_element(By.XPATH, "//button[@aria-label='Menu actions trigger']") dropdown_trigger.click() logging.info("Dropdown trigger clicked.") action = ActionChains(driver) logging.info("Attempting to find the dropdown item for download.") dropdown_item = driver.find_element(By.XPATH, "//div[@title='Download']") action.move_to_element(dropdown_item).perform() logging.info("The dropdown item was found.") # 点击导出PDF按钮 logging.info("Attempting to click the 'Export to PDF' button.") export_to_pdf_button = WebDriverWait(driver, 3).until( EC.element_to_be_clickable((By.XPATH, "//div[@role='button'][contains(text(), 'Export to PDF')]")) ) export_to_pdf_button.click() logging.info("Export to PDF button clicked, waiting for file to download.") # 优化后的下载等待逻辑 start_time = time.time() download_completed = False while time.time() - start_time < 120: # 延长至120秒等待 all_files = os.listdir(unique_dir) pdf_files = [f for f in all_files if f.endswith(".pdf")] temp_files = [f for f in all_files if f.endswith(".part")] if pdf_files: download_completed = True break elif temp_files: # 有临时文件,说明正在下载,等待2秒再检查 time.sleep(2) else: time.sleep(1) if not download_completed: raise Exception("File download timed out.") finally: driver.quit() logging.info("Driver quit.")
内容的提问来源于stack exchange,提问作者Gonzalo Ramos Farinho
相关产品推荐
相关产品推荐

