Python3实现文件夹遍历及相似文件夹检测的问题求助
相似文件夹查找工具:相似度对比实现思路
第一步:重构文件夹数据存储
你当前的代码收集的是完整文件夹路径,首先需要把文件夹名称和路径关联起来,方便后续对比和结果展示。可以用字典存储,键是完整路径,值是对应的文件夹名称:
import os from difflib import SequenceMatcher path = input("Where you want to look?") folder_dict = {} print("Here's your folders:") for dirname in os.listdir(path): full_path = os.path.join(path, dirname) if os.path.isdir(full_path): folder_name = os.path.basename(full_path) folder_dict[full_path] = folder_name print(full_path) print(f"\n{len(folder_dict)} folders found!")
第二步:实现相似度对比逻辑
针对文件名不完全相同的场景,推荐两种实用的相似度计算方式:
方式1:基于序列匹配的相似度(推荐)
用Python标准库difflib.SequenceMatcher计算字符串的相似度,返回0-1的数值,越接近1表示越相似,该方法会考虑字符顺序,更符合文件名相似的实际场景(比如"Project_v1"和"Project_v2"会被判定为高度相似):
def calculate_similarity(str1, str2): return SequenceMatcher(None, str1, str2).ratio()
方式2:基于字符集合的Jaccard相似度(按你的拆字符思路实现)
把文件名拆成字符集合,通过计算交集与并集的比例得到相似度,适合关注字符组成而非顺序的场景:
def jaccard_similarity(str1, str2): set1 = set(str1) set2 = set(str2) intersection = len(set1 & set2) union = len(set1 | set2) return intersection / union if union != 0 else 0
第三步:遍历所有文件夹对找出相似项
设置一个相似度阈值(比如0.7,可根据需求调整),遍历所有文件夹两两组合,筛选出相似度达标的结果:
# 可自定义相似度阈值 similarity_threshold = 0.7 print("\n相似文件夹对:") folder_paths = list(folder_dict.keys()) # 两两对比,避免重复检查 for i in range(len(folder_paths)): for j in range(i + 1, len(folder_paths)): path_a = folder_paths[i] path_b = folder_paths[j] name_a = folder_dict[path_a] name_b = folder_dict[path_b] # 选择你想用的相似度计算函数 sim_score = calculate_similarity(name_a, name_b) # sim_score = jaccard_similarity(name_a, name_b) if sim_score >= similarity_threshold: print(f"相似度 {sim_score:.2f}:") print(f" - {path_a}") print(f" - {path_b}\n")
补充说明
- 如果需要递归遍历子文件夹,可改用
os.walk()替代os.listdir() - 相似度阈值可根据实际需求调整,对相似度要求高就设0.8,要求宽松就设0.6
内容的提问来源于stack exchange,提问作者dotashope
相关产品推荐
相关产品推荐

