You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的re模块按文件名模式分组归类文件?

按文件名模式分组文件的解决方案

你可以用字典存储不同分组的文件路径,结合正则匹配或字符串判断来实现归类。这里提供两种实用的实现方式:

方法一:正则精准匹配

这种方式能准确提取文件名中usc_后的三位模式(hhx、hhh、hhz),避免误匹配:

import os
import re

# 初始化分组字典,键为目标模式,值为对应文件路径列表
file_groups = {
    'hhx': [],
    'hhh': [],
    'hhz': []
}

# 正则匹配usc_后接三位模式+数字的格式
pattern = re.compile(r'usc_(hh[xyz])\d+')

folder_path = "ginsburgdata/USC1/USC1"
for file in os.listdir(folder_path):
    # 统一转小写,兼容文件名大小写不一致的情况
    match = pattern.match(file.lower())
    if match:
        group_key = match.group(1)
        if group_key in file_groups:
            full_path = os.path.join(os.getcwd(), folder_path, file)
            file_groups[group_key].append(full_path)

# 输出分组结果
for group, files in file_groups.items():
    print(f"分组 {group}:")
    for f in files:
        print(f"  {f}")

方法二:字符串包含判断

如果文件名格式固定,直接用字符串包含判断更简洁:

import os

file_groups = {
    'hhx': [],
    'hhh': [],
    'hhz': []
}

folder_path = "ginsburgdata/USC1/USC1"
for file in os.listdir(folder_path):
    file_lower = file.lower()
    full_path = os.path.join(os.getcwd(), folder_path, file)
    if 'hhx' in file_lower:
        file_groups['hhx'].append(full_path)
    elif 'hhh' in file_lower:
        file_groups['hhh'].append(full_path)
    elif 'hhz' in file_lower:
        file_groups['hhz'].append(full_path)

# 查看分组结果
print(file_groups)

关键说明

  • 用字典存储分组比多个独立列表更清晰,可直接通过模式键快速访问对应文件
  • 统一转小写是为了避免文件名大小写差异导致的匹配失败(比如注释里的HHZ和实际文件名的hhz)
  • 用变量存储文件夹路径,避免重复书写路径字符串出错

内容的提问来源于stack exchange,提问作者danny_mai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 17:15:27