You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的Selenium(Firefox)确保文件完整下载并解决异常

问题

我是一名做Python项目的学生,需要通过Firefox下载指定网页上的CSV文件(文件名为prix-des-carburants-en-france-flux-instantane-v2,大小13.8Mo)。我写了脚本模拟人工操作:访问网站→获取文件名和更新时间→进入“Export”板块→点击CSV下载链接。

但偶尔(约1%概率)会出现两种下载中断问题:

  • 文件显示下载完成,但大小不足13.8Mo,可能是任意小于该值的大小;
  • 浏览器弹出弹窗:“Cancel all downloads? If you exit now, a download in progress will be canceled. Do you really want to leave?”,无人干预就无法继续。

我不想用os模块或requests方案,更倾向于保留当前模拟人工点击的方式,用Firefox自带下载系统,平时运行都正常,也不需要try-except。

我试过用wait_for_fully_downloaded方法监控文件大小是否停止增长来判断下载完成,但问题依旧。现在想请教:

  1. 能不能用with语句解决第二种弹窗(浏览器在下载完成前关闭)的问题?
  2. 是否应该用quit()?quit()会在执行崩溃时关闭浏览器吗?
  3. 第一种问题,增大check_interval参数可行吗?但这样会导致程序空转。

以下是我的代码:

"""
    Module which provides methods for scraping data using Firefox webdriver.
"""
import time
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import WebDriverException


class FirefoxScraperHolder:
    """
        Class for scraping data using Firefox webdriver.
    """
    def __init__(self, target_url):
        """
            Initialize a FirefoxScraperHolder instance.

            :param target_url: The URL to scrape data from.
        """
        self.cwf = Path(__file__).resolve().parent

        self.options = webdriver.FirefoxOptions()
        self.driver = webdriver.Firefox(options=self.set_preferences())
        self.target_url = target_url

        self._updated_data_date = None
        self._csv_id = None

    def set_preferences(self):
        """
            Set Firefox webdriver preferences
            (Here only for downloading files)

            :return: The configured Firefox options.
        """
        # 2 means we use chosen directory as download folder
        self.options.set_preference("browser.download.folderList", 2)
        self.options.set_preference("browser.download.dir", str(self.cwf))
        return self.options

    @property
    def updated_data_date(self):
        """
        Get the last updated data date.

        :return: The last updated data date.
        :rtype: str
        """
        return self._updated_data_date

    @property
    def csv_id(self):
        """
        Get the filename of the downloaded CSV.

        :return: The filename of the downloaded CSV.
        :rtype: str
        """
        return self._csv_id

    def perform_scraping(self, aria_label, ng_if):
        """
            Perform the scraping process.

            :param aria_label: ARIA label for the CSV element.
            :param ng_if: NG-if attribute for updated data date element.
        """
        try:
            with self.driver:
                self.driver.maximize_window()
                self.driver.get(self.target_url)

                # Retrieves csv information
                self.click_on(By.LINK_TEXT, "Informations")
                self._updated_data_date = self.retrieve_text_info(
                    By.CSS_SELECTOR,
                    f"[ng-if='{ng_if}']")
                self._csv_id = self.retrieve_text_info(
                    By.CLASS_NAME,
                    'ods-dataset-metadata-block__metadata-value'
                ) + '.csv'

                # Download csv
                self.click_on(By.LINK_TEXT, "Export")
                self.click_on(By.CSS_SELECTOR, f"[aria-label='{aria_label}']")
                self.wait_for_fully_downloaded()

        except WebDriverException as exception:
            print(f"An error occurred during the get operation: {exception}")

    def click_on(self, find_by, value):
        """
            Click on a web element identified by 'find_by' and 'value'.

            :param find_by: The method used to find the element
            (e.g., By.LINK_TEXT).
            :param value: The value to search for.
        """
        # Here 'wait' and 'EC' avoid error due to the loading of the website
        wait = WebDriverWait(self.driver, 20)
        element = wait.until(EC.element_to_be_clickable((find_by, value)))
        element.click()

    def remove_cwf_existing_csvs(self):
        """
            Remove existing CSV files from the current working folder.
        """
        for file in self.cwf.glob('*.csv'):
            file.unlink(missing_ok=True)

    def retrieve_text_info(self, find_by, value):
        """
            Retrieve text information of a web element identified
            by 'find_by' and 'value'.

            :param find_by: The method used to find the element
            (e.g., By.CSS_SELECTOR).
            :param value: The value to search for.
            :return: The text information of the web element.
            :rtype: str
        """
        # Here 'wait' and 'EC' avoid error due to the loading of the website
        wait = WebDriverWait(self.driver, 20)
        info = wait.until(EC.visibility_of_element_located((find_by, value)))
        return info.text

    def wait_for_fully_downloaded(self, timeout=60, check_interval=1):
        """
            Wait for a file to be fully downloaded.

            :param timeout: Maximum time to wait in seconds.
            Default is 60 seconds.
            :param check_interval: Interval for checking file size
             in seconds. Default is 1 second.
        """
        file_path = self.cwf / self._csv_id
        start_time = time.time()

        while time.time() - start_time < timeout:
            if file_path.is_file():
                initial_size = file_path.stat().st_size
                time.sleep(check_interval)
                # Checks if the file size changes during check_interval
                if file_path.stat().st_size == initial_size:
                    return
        return

# Here my main code:

# target_url = ('https://data.economie.gouv.fr/explore/dataset/prix-des-' 
#               'carburants-en-france-flux-instantane-v2/')
# csv_aria_label = 'Dataset export (CSV)'
# updated_data_date_ng_if = 'ctx.dataset.metas.data_processed'
# 
# # Retrieves datas
# firefox_scraper = FirefoxScraperHolder(target_url)
# firefox_scraper.remove_cwf_existing_csvs()
# firefox_scraper.perform_scraping(csv_aria_label, updated_data_date_ng_if)

解决方案与答疑

1. with语句与弹窗问题

你已经在用with self.driver,但弹窗出现的核心原因是**wait_for_fully_downloaded误判下载完成,导致with代码块提前结束,触发driver.quit()时下载还在进行**。

要解决这个问题,需要优化下载完成的判断逻辑,确保只有下载彻底完成后,才让with代码块结束:

  • 利用Firefox的临时文件机制:下载时会生成带.part后缀的临时文件,完成后会自动移除该后缀;
  • 加入文件大小校验,和预期的13.8Mo(≈14460928字节)做比对,允许1%左右的误差。

修改后的wait_for_fully_downloaded示例:

def wait_for_fully_downloaded(self, timeout=60, check_interval=1, expected_size_bytes=14460928):
    file_path = self.cwf / self._csv_id
    temp_file_path = file_path.with_suffix('.part')
    start_time = time.time()

    while time.time() - start_time < timeout:
        # 先确认临时文件消失,且目标文件存在
        if not temp_file_path.exists() and file_path.exists():
            current_size = file_path.stat().st_size
            # 校验大小误差在允许范围内
            if abs(current_size - expected_size_bytes) / expected_size_bytes < 0.01:
                return
            # 大小不符则继续等待,可能还在收尾
            time.sleep(check_interval)
        else:
            time.sleep(check_interval)
    
    # 超时后给出提示
    if file_path.exists():
        print(f"下载文件大小异常:{file_path.stat().st_size}字节(预期{expected_size_bytes}字节)")
    else:
        print("下载超时未完成")

同时可以给Firefox添加偏好设置,禁止下载相关弹窗:

def set_preferences(self):
    self.options.set_preference("browser.download.folderList", 2)
    self.options.set_preference("browser.download.dir", str(self.cwf))
    # 关闭下载弹窗相关设置
    self.options.set_preference("browser.download.manager.showWhenStarting", False)
    self.options.set_preference("browser.download.manager.focusWhenStarting", False)
    self.options.set_preference("browser.helperApps.neverAsk.saveToDisk", "text/csv")
    return self.options

2. quit()的使用说明

  • 是否应该用quit()?:是的。quit()会彻底关闭浏览器进程和所有窗口,释放系统资源,你用with语句自动调用quit()是正确的做法。
  • quit()会在执行崩溃时关闭浏览器吗?:大部分情况下会。with语句的上下文管理器会在代码块结束(包括异常抛出时)自动调用driver.quit(),但如果是致命崩溃(比如Python解释器退出),可能无法执行。可以给类添加兜底的销毁方法:
def __del__(self):
    try:
        self.driver.quit()
    except Exception:
        pass

3. check_interval与文件大小不足问题

增大check_interval只能减少空转次数,解决不了文件大小不足的根本问题。该问题的常见原因是网络波动中断下载,或原有的等待逻辑误判下载完成。

优化思路:

  • 结合临时文件和大小校验判断下载完成,避免误判;
  • 添加重试机制,检测到文件大小异常时重新触发下载;
  • 不要单纯依赖“大小停止增长”,因为网络卡顿可能导致大小短暂停滞。

示例:在perform_scraping中加入重试逻辑

def perform_scraping(self, aria_label, ng_if, max_retries=2):
    retry_count = 0
    while retry_count <= max_retries:
        try:
            with self.driver:
                self.driver.maximize_window()
                self.driver.get(self.target_url)

                # 获取文件信息
                self.click_on(By.LINK_TEXT, "Informations")
                self._updated_data_date = self.retrieve_text_info(
                    By.CSS_SELECTOR, f"[ng-if='{ng_if}']")
                self._csv_id = self.retrieve_text_info(
                    By.CLASS_NAME, 'ods-dataset-metadata-block__metadata-value'
                ) + '.csv'

                # 触发下载并等待完成
                self.click_on(By.LINK_TEXT, "Export")
                self.click_on(By.CSS_SELECTOR, f"[aria-label='{aria_label}']")
                self.wait_for_fully_downloaded()

                # 最终校验大小
                file_path = self.cwf / self._csv_id
                expected_size = 14460928
                if abs(file_path.stat().st_size - expected_size) / expected_size > 0.01:
                    raise Exception(f"文件大小不符合预期")
                
                print("下载成功")
                return
        except (WebDriverException, Exception) as e:
            print(f"第{retry_count+1}次尝试失败:{e}")
            retry_count += 1
            # 重新初始化浏览器
            self.driver = webdriver.Firefox(options=self.set_preferences())
    print(f"已重试{max_retries}次,下载失败")

内容的提问来源于stack exchange,提问作者Hugocrt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 12:37:04