如何用Python的re模块按文件名模式分组归类文件?
按文件名模式分组文件的解决方案
你可以用字典存储不同分组的文件路径,结合正则匹配或字符串判断来实现归类。这里提供两种实用的实现方式:
方法一:正则精准匹配
这种方式能准确提取文件名中usc_后的三位模式(hhx、hhh、hhz),避免误匹配:
import os import re # 初始化分组字典,键为目标模式,值为对应文件路径列表 file_groups = { 'hhx': [], 'hhh': [], 'hhz': [] } # 正则匹配usc_后接三位模式+数字的格式 pattern = re.compile(r'usc_(hh[xyz])\d+') folder_path = "ginsburgdata/USC1/USC1" for file in os.listdir(folder_path): # 统一转小写,兼容文件名大小写不一致的情况 match = pattern.match(file.lower()) if match: group_key = match.group(1) if group_key in file_groups: full_path = os.path.join(os.getcwd(), folder_path, file) file_groups[group_key].append(full_path) # 输出分组结果 for group, files in file_groups.items(): print(f"分组 {group}:") for f in files: print(f" {f}")
方法二:字符串包含判断
如果文件名格式固定,直接用字符串包含判断更简洁:
import os file_groups = { 'hhx': [], 'hhh': [], 'hhz': [] } folder_path = "ginsburgdata/USC1/USC1" for file in os.listdir(folder_path): file_lower = file.lower() full_path = os.path.join(os.getcwd(), folder_path, file) if 'hhx' in file_lower: file_groups['hhx'].append(full_path) elif 'hhh' in file_lower: file_groups['hhh'].append(full_path) elif 'hhz' in file_lower: file_groups['hhz'].append(full_path) # 查看分组结果 print(file_groups)
关键说明
- 用字典存储分组比多个独立列表更清晰,可直接通过模式键快速访问对应文件
- 统一转小写是为了避免文件名大小写差异导致的匹配失败(比如注释里的
HHZ和实际文件名的hhz) - 用变量存储文件夹路径,避免重复书写路径字符串出错
内容的提问来源于stack exchange,提问作者danny_mai
相关产品推荐
相关产品推荐

