You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS EC2中Python Selenium+Firefox下载PDF文件失败问题

AWS EC2上Selenium Firefox无法保存PDF到指定目录的问题排查与解决

问题背景

开发Python项目时,通过Selenium+Firefox自动从网站下载PDF并上传至存储桶:

  • 开发环境基于Docker容器,搭配Selenium 3.141.0/3.8.0/4.18.1版本与Firefox驱动,PDF可正常保存至/tmp目录
  • 部署到AWS EC2实例后,应用与网站交互全程无报错,指定的unique_dir目录已生成,机器剩余21GB可用空间,但目录始终为空,PDF文件无处可寻

可能原因

  • Firefox无头模式限制:EC2环境下,无头模式的Firefox可能因沙箱机制或/dev/shm资源限制,导致下载路径配置不生效
  • MIME类型不匹配:网站返回的PDF文件Content-Type可能不是application/pdf,Firefox未触发自动保存逻辑(无头模式下无法弹出保存对话框)
  • 版本兼容性问题:EC2上的Firefox或Geckodriver版本与开发环境不一致,导致下载偏好配置失效
  • 等待逻辑缺陷:原逻辑仅检查.pdf后缀文件,但下载过程中会生成临时文件(如.part),可能因等待时间不足或未识别临时文件导致超时

解决方案

1. 调整Firefox无头模式配置

添加适配EC2环境的参数,补充下载相关偏好:

  • --no-sandbox:禁用沙箱,避免权限限制
  • --disable-dev-shm-usage:绕过/dev/shm的大小限制
  • 强制指定下载目录、禁用内置PDF查看器,避免在线打开PDF

2. 扩展MIME类型覆盖范围

将常见的PDF相关MIME类型都加入自动保存列表,覆盖application/octet-stream、application/x-pdf等可能的类型

3. 优化下载等待逻辑

增加对临时下载文件(.part)的检测,延长等待时间,确保文件完成下载后再退出循环

4. 排查日志与环境一致性

  • 查看/tmp/geckodriver.log中的下载相关日志,定位权限或路径错误
  • 确保EC2上的Firefox、Geckodriver版本与开发环境Docker内完全一致

修改后的代码示例

import os
import stat
import uuid
import time
import logging
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.firefox.options import Options
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

login_url = 'https://my-web.com/login/'
dashboard_url = f'https://my-web.com/pdf-button-download-view'
unique_dir = os.path.join("/tmp/pdfs", str(uuid.uuid4()))
os.makedirs(unique_dir, exist_ok=True)

# 确保目录权限足够
os.chmod(unique_dir, stat.S_IRWXU | stat.S_IRWXG | stat.S_IRWXO)

options = Options()
options.headless = True
# 适配EC2无头环境的关键参数
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")
options.add_argument("--disable-gpu")
log_path = "/tmp/geckodriver.log"

# 配置Firefox下载偏好
profile = webdriver.FirefoxProfile()
profile.set_preference("browser.download.folderList", 2)
profile.set_preference("browser.download.manager.showWhenStarting", False)
profile.set_preference("browser.download.dir", unique_dir)
# 覆盖所有可能的PDF MIME类型
profile.set_preference("browser.helperApps.neverAsk.saveToDisk", 
                      "application/pdf,application/octet-stream,application/x-pdf")
# 强制使用指定下载目录
profile.set_preference("browser.download.useDownloadDir", True)
# 禁用内置PDF查看器,强制下载
profile.set_preference("pdfjs.disabled", True)
# 关闭下载管理器相关弹窗(无头模式无法处理)
profile.set_preference("browser.download.manager.useWindow", False)
profile.set_preference("browser.download.manager.focusWhenStarting", False)
profile.set_preference("browser.download.manager.closeWhenDone", True)

driver = webdriver.Firefox(options=options, firefox_profile=profile, service_log_path=log_path)

try:
    driver.get(login_url)
    time.sleep(10)

    # 登录流程
    driver.find_element(By.ID, "username").send_keys("admin")
    driver.find_element(By.ID, "password").send_keys("admin")
    driver.find_element(By.CSS_SELECTOR, "form").submit()
    time.sleep(10)

    # 进入仪表盘
    driver.get(dashboard_url)
    time.sleep(10)

    logging.info(f"Dashboard URL loaded: {driver.current_url}")

    # 触发下拉菜单
    dropdown_trigger = driver.find_element(By.XPATH, "//button[@aria-label='Menu actions trigger']")
    dropdown_trigger.click()
    logging.info("Dropdown trigger clicked.")

    action = ActionChains(driver)
    logging.info("Attempting to find the dropdown item for download.")

    dropdown_item = driver.find_element(By.XPATH, "//div[@title='Download']")
    action.move_to_element(dropdown_item).perform()
    logging.info("The dropdown item was found.")

    # 点击导出PDF按钮
    logging.info("Attempting to click the 'Export to PDF' button.")
    export_to_pdf_button = WebDriverWait(driver, 3).until(
        EC.element_to_be_clickable((By.XPATH, "//div[@role='button'][contains(text(), 'Export to PDF')]"))
    )
    export_to_pdf_button.click()
    logging.info("Export to PDF button clicked, waiting for file to download.")

    # 优化后的下载等待逻辑
    start_time = time.time()
    download_completed = False
    while time.time() - start_time < 120:  # 延长至120秒等待
        all_files = os.listdir(unique_dir)
        pdf_files = [f for f in all_files if f.endswith(".pdf")]
        temp_files = [f for f in all_files if f.endswith(".part")]
        
        if pdf_files:
            download_completed = True
            break
        elif temp_files:
            # 有临时文件,说明正在下载,等待2秒再检查
            time.sleep(2)
        else:
            time.sleep(1)
    
    if not download_completed:
        raise Exception("File download timed out.")

finally:
    driver.quit()
    logging.info("Driver quit.")

内容的提问来源于stack exchange,提问作者Gonzalo Ramos Farinho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 18:10:58