Python:如何可靠检测下载文件是否完成?
可靠检测下载文件完成的Python方案
针对大文件下载未完成就被移动导致文件损坏的问题,以下是几种实用可靠的解决思路:
1. 检测文件是否处于可安全操作状态(核心方案)
文件正在被下载工具写入时,通常会被锁定,无法以独占模式打开。通过循环尝试打开文件的方式,可准确判断下载是否完成:
import os import time import shutil from watchdog.events import FileSystemEventHandler import threading def is_file_ready(file_path): """判断文件是否已完成写入,可安全操作""" if not os.path.exists(file_path): return False # 尝试以只读模式打开文件,成功则说明未被占用 try: with open(file_path, 'rb') as f: return True except (PermissionError, OSError): return False def move_file(src, dst): # 循环等待文件可访问 while not is_file_ready(src): time.sleep(0.5) # 额外校验文件大小稳定,应对缓冲写入场景 prev_size = -1 while True: current_size = os.path.getsize(src) if current_size == prev_size: break prev_size = current_size time.sleep(0.5) shutil.move(src, dst) class DownloadHandler(FileSystemEventHandler): def on_created(self, event): if event.src_path.endswith(".zip") and not event.is_directory: threading.Thread(target=move_file, args=(event.src_path, dst)).start()
2. 利用浏览器下载的临时后缀特征
多数浏览器下载时会使用临时文件名:
- Chrome/Edge:
.crdownload后缀 - Firefox:
.part后缀 - Safari:下载过程中文件存于临时目录,完成后才移至下载文件夹
修改监控逻辑,只在文件后缀转为目标格式(如.zip)时触发移动:
class DownloadHandler(FileSystemEventHandler): def on_modified(self, event): # 监控文件后缀变化,临时文件转为zip时处理 if event.src_path.endswith(".zip") and not event.is_directory: threading.Thread(target=move_file, args=(event.src_path, dst)).start() def on_created(self, event): # 直接忽略临时后缀的文件 if any(event.src_path.endswith(ext) for ext in [".crdownload", ".part"]): return
3. 结合文件修改时间判断
等待文件修改时间不再变化超过指定时长(如2秒),适用于不会锁定文件的下载工具:
def wait_for_file_stable(file_path, wait_duration=2): if not os.path.exists(file_path): return False last_modified = os.path.getmtime(file_path) while True: current_modified = os.path.getmtime(file_path) if current_modified == last_modified: time.sleep(wait_duration) # 二次确认修改时间未变化 if os.path.getmtime(file_path) == current_modified: return True last_modified = current_modified time.sleep(0.5) # 在move_file中调用 def move_file(src, dst): while not wait_for_file_stable(src): pass shutil.move(src, dst)
注意事项
- 优先结合文件占用检测和临时后缀判断,这两种方式结合能覆盖绝大多数浏览器和下载工具的场景。
- 避免使用固定时长的
time.sleep(),循环检测更适配不确定时长的大文件下载。
内容的提问来源于stack exchange,提问作者Octa
相关产品推荐
相关产品推荐

