如何在驱动器一级文件夹中批量运行Python去重脚本?
批量处理网络驱动器一级文件夹内重复文件
问题背景
我正在处理网络驱动器的数据清理工作,该驱动器包含1000多个一级文件夹,每个文件夹下还有若干子文件夹。现有一份脚本可清理单个文件夹内的.txt、.bmp格式重复文件,但需要手动逐个选择文件夹,效率极低。直接将脚本应用于整个驱动器会误删不同一级文件夹下的同名文件(例如Z:/Folder1和Z:/Folder2中的text.txt需各自保留一份),不符合需求。需要修改脚本实现自动遍历所有一级文件夹并分别执行去重操作。
原脚本如下:
from tkinter.filedialog import askdirectory # Importing required libraries. from tkinter import Tk import os import hashlib from pathlib import Path # We don't want the GUI window of # tkinter to be appearing on our screen Tk().withdraw() # Dialog box for selecting a folder. file_path = askdirectory(title="Select a folder") # Listing out all the files # inside our root folder. list_of_files = os.walk(file_path) # In order to detect the duplicate # files we are going to define an empty dictionary. unique_files = dict() for root, folders, files in list_of_files: # Running a for loop on all the files for file in files: # Finding complete file path file_path = Path(os.path.join(root, file)) # Converting all the content of # our file into md5 hash. Hash_file = hashlib.md5(open(file_path, 'rb').read()).hexdigest() # If file hash has already # # been added we'll simply delete that file if Hash_file not in unique_files: unique_files[Hash_file] = file_path else: if file.endswith((".txt",".bmp")): os.remove(file_path) print(f"{file_path} has been deleted")
修改后的脚本
from tkinter.filedialog import askdirectory from tkinter import Tk import os import hashlib from pathlib import Path Tk().withdraw() # 选择根驱动器(比如Z盘) root_drive = askdirectory(title="选择网络驱动器根目录") # 遍历根目录下的所有一级文件夹 for entry in os.scandir(root_drive): if entry.is_dir(): # 每个一级文件夹单独初始化去重字典 unique_files = dict() folder_path = entry.path print(f"开始处理文件夹: {folder_path}") # 遍历当前一级文件夹及其子目录 for root, folders, files in os.walk(folder_path): for file in files: file_path = Path(os.path.join(root, file)) # 只处理指定格式文件 if not file.endswith((".txt", ".bmp")): continue try: # 计算文件MD5哈希 Hash_file = hashlib.md5(open(file_path, 'rb').read()).hexdigest() if Hash_file not in unique_files: unique_files[Hash_file] = file_path else: os.remove(file_path) print(f"已删除重复文件: {file_path}") except Exception as e: print(f"处理文件 {file_path} 时出错: {str(e)}") print(f"文件夹 {folder_path} 处理完成\n")
关键修改说明
- 替换单文件夹选择为根目录选择:让用户选择整个网络驱动器根目录(如Z盘),而非单个文件夹。
- 遍历所有一级文件夹:使用
os.scandir遍历根目录下的所有一级目录,跳过文件项。 - 每个文件夹独立去重:针对每个一级文件夹,重新初始化
unique_files字典,确保不同一级文件夹的文件不会互相判定为重复。 - 提前过滤文件格式:在计算哈希前先判断文件格式,减少不必要的哈希计算,提升效率。
- 增加异常处理:捕获文件读取或删除时的异常,避免脚本因单个文件问题中断。
内容的提问来源于stack exchange,提问作者Paul W
相关产品推荐
相关产品推荐

