You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在驱动器一级文件夹中批量运行Python去重脚本?

批量处理网络驱动器一级文件夹内重复文件

问题背景

我正在处理网络驱动器的数据清理工作,该驱动器包含1000多个一级文件夹,每个文件夹下还有若干子文件夹。现有一份脚本可清理单个文件夹内的.txt、.bmp格式重复文件,但需要手动逐个选择文件夹,效率极低。直接将脚本应用于整个驱动器会误删不同一级文件夹下的同名文件(例如Z:/Folder1和Z:/Folder2中的text.txt需各自保留一份),不符合需求。需要修改脚本实现自动遍历所有一级文件夹并分别执行去重操作。

原脚本如下:

from tkinter.filedialog import askdirectory

# Importing required libraries.
from tkinter import Tk
import os
import hashlib
from pathlib import Path

# We don't want the GUI window of
# tkinter to be appearing on our screen
Tk().withdraw()

# Dialog box for selecting a folder.
file_path = askdirectory(title="Select a folder")

# Listing out all the files
# inside our root folder.
list_of_files = os.walk(file_path)

# In order to detect the duplicate
# files we are going to define an empty dictionary.
unique_files = dict()

for root, folders, files in list_of_files:

    # Running a for loop on all the files
    for file in files:

        # Finding complete file path
        file_path = Path(os.path.join(root, file))

        # Converting all the content of
        # our file into md5 hash.
        Hash_file = hashlib.md5(open(file_path, 'rb').read()).hexdigest()

        # If file hash has already #
        # been added we'll simply delete that file
        if Hash_file not in unique_files:
            unique_files[Hash_file] = file_path
        else:
            if file.endswith((".txt",".bmp")):
                os.remove(file_path)
                print(f"{file_path} has been deleted")

修改后的脚本

from tkinter.filedialog import askdirectory
from tkinter import Tk
import os
import hashlib
from pathlib import Path

Tk().withdraw()
# 选择根驱动器(比如Z盘)
root_drive = askdirectory(title="选择网络驱动器根目录")

# 遍历根目录下的所有一级文件夹
for entry in os.scandir(root_drive):
    if entry.is_dir():
        # 每个一级文件夹单独初始化去重字典
        unique_files = dict()
        folder_path = entry.path
        print(f"开始处理文件夹: {folder_path}")
        
        # 遍历当前一级文件夹及其子目录
        for root, folders, files in os.walk(folder_path):
            for file in files:
                file_path = Path(os.path.join(root, file))
                # 只处理指定格式文件
                if not file.endswith((".txt", ".bmp")):
                    continue
                
                try:
                    # 计算文件MD5哈希
                    Hash_file = hashlib.md5(open(file_path, 'rb').read()).hexdigest()
                    
                    if Hash_file not in unique_files:
                        unique_files[Hash_file] = file_path
                    else:
                        os.remove(file_path)
                        print(f"已删除重复文件: {file_path}")
                except Exception as e:
                    print(f"处理文件 {file_path} 时出错: {str(e)}")
        print(f"文件夹 {folder_path} 处理完成\n")

关键修改说明

  • 替换单文件夹选择为根目录选择:让用户选择整个网络驱动器根目录(如Z盘),而非单个文件夹。
  • 遍历所有一级文件夹:使用os.scandir遍历根目录下的所有一级目录,跳过文件项。
  • 每个文件夹独立去重:针对每个一级文件夹,重新初始化unique_files字典,确保不同一级文件夹的文件不会互相判定为重复。
  • 提前过滤文件格式:在计算哈希前先判断文件格式,减少不必要的哈希计算,提升效率。
  • 增加异常处理:捕获文件读取或删除时的异常,避免脚本因单个文件问题中断。

内容的提问来源于stack exchange,提问作者Paul W

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 13:19:00