You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现文件名到对应类型的精准分类,避免相似标识误匹配?

文件名分类匹配问题

场景说明

我有一个包含大量目录文件名的列表,示例如下:

file_list = ['apple-20220103.csv', 'apple_tea-20220304.csv', '20220203-apple_town.csv', 'apple_town20220101.csv']

文件类型存储在一个.csv文件中,内容如下:

,type
0,apple
1,apple_tea
2,apple_town

目标是将列表中的每个文件名分类到对应的类型中并存入字典,最终预期结果:

result_dict = {
     'apple':['apple-20220103.csv'],
     'apple_tea':['apple_tea-20220304.csv'],
     'apple_town':['20220203-apple_town.csv', 'apple_town20220101.csv']
}

核心问题

使用简单正则匹配时,apple类型会误匹配到包含apple字样的apple_tea、apple_town类文件,如何确保精准匹配?


解决方案

方法1:按类型长度倒序匹配

核心逻辑是优先匹配更长的类型字符串,长类型(如apple_tea)是短类型(apple)的扩展,先匹配长类型就能避免短类型误抢匹配。

步骤:

  1. 从CSV读取类型列表,按类型字符串长度从长到短排序
  2. 遍历每个文件名,依次尝试匹配排序后的类型,匹配成功就归入对应类别并跳出循环,避免重复匹配

示例代码:

import csv

# 读取类型列表
types = []
with open('types.csv', 'r', newline='') as f:
    reader = csv.DictReader(f)
    for row in reader:
        types.append(row['type'])

# 按类型长度倒序排序,长类型优先匹配
types_sorted = sorted(types, key=lambda x: len(x), reverse=True)

file_list = ['apple-20220103.csv', 'apple_tea-20220304.csv', '20220203-apple_town.csv', 'apple_town20220101.csv']
result_dict = {t: [] for t in types}

for filename in file_list:
    for t in types_sorted:
        if t in filename:
            result_dict[t].append(filename)
            break  # 找到匹配类型后跳出,避免短类型重复匹配

print(result_dict)

方法2:正则边界精准匹配

通过正则的上下文限制,确保apple匹配的是独立标识,而非其他类型的前缀部分。

示例代码:

import re
import csv

# 读取类型
types = []
with open('types.csv', 'r', newline='') as f:
    reader = csv.DictReader(f)
    for row in reader:
        types.append(row['type'])

file_list = ['apple-20220103.csv', 'apple_tea-20220304.csv', '20220203-apple_town.csv', 'apple_town20220101.csv']
result_dict = {t: [] for t in types}

# 为每个类型构建精准匹配的正则模式
pattern_dict = {
    'apple': re.compile(r'\bapple(?=-|\.csv|$)'),  # 匹配apple后紧跟连字符、.csv或结尾
    'apple_tea': re.compile(r'apple_tea'),
    'apple_town': re.compile(r'apple_town')
}

for filename in file_list:
    for t in types:
        if pattern_dict[t].search(filename):
            result_dict[t].append(filename)
            break

print(result_dict)

方法3:提取文件名标识部分匹配

如果文件名格式有规律,直接提取可能的标识片段再与类型列表对比,避免子串误匹配。

示例代码:

import csv

types = []
with open('types.csv', 'r', newline='') as f:
    reader = csv.DictReader(f)
    for row in reader:
        types.append(row['type'])

file_list = ['apple-20220103.csv', 'apple_tea-20220304.csv', '20220203-apple_town.csv', 'apple_town20220101.csv']
result_dict = {t: [] for t in types}
type_set = set(types)
types_sorted = sorted(types, key=lambda x: len(x), reverse=True)

for filename in file_list:
    # 去掉文件后缀
    name_part = filename.replace('.csv', '')
    candidates = []
    # 按连字符分割提取候选
    candidates.extend(name_part.split('-'))
    # 提取开头/结尾匹配的类型候选
    for t in types_sorted:
        if name_part.startswith(t) or name_part.endswith(t):
            candidates.append(t)
    # 查找匹配的类型
    for cand in set(candidates):
        if cand in type_set:
            result_dict[cand].append(filename)
            break

print(result_dict)

内容的提问来源于stack exchange,提问作者Helmi Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 01:40:44