Python 2.6中如何排除文件名包含特定子串的文件?
解决方法:排除包含特定子串的匹配文件
你的问题出在fnmatch使用的是shell通配符语法,不是标准正则表达式,所以像(?!uat)这种正则负向预查的写法在fnmatch里完全不生效,而[^uat]只是排除单个字符(u、a、t中的任意一个),不是排除整个"uat"子串,这就是为什么你的尝试都没得到预期结果。
下面给你两种可行的修改方案:
方案1:先匹配基础规则,再过滤排除子串(最简单直接)
保持原有的fnmatch匹配逻辑,在循环里额外添加判断,排除包含"uat"或"test"的文件:
#!/usr/bin/env python import xml.etree.ElementTree as ET, sys, os, re, fnmatch param = sys.argv[1] client = param.split('_')[0] market = param.split('_')[1] suffix = param.split('_')[2] toapex_pattern = market + '*2apex*' + client + '*' + '.xml' files_dir = '/some/dir' config_files = os.listdir(files_dir) for f in config_files: if fnmatch.fnmatch(f, toapex_pattern): # 新增:排除包含uat或test的文件 if 'uat' not in f and 'test' not in f: print(f)
运行这个修改后的脚本,就能得到你想要的fiot_csv2apex_nomura.xml。
方案2:改用正则表达式匹配(更灵活)
如果你想用正则来实现完整的匹配+排除逻辑,需要放弃fnmatch,直接使用Python的re模块。注意正则里的*和shell通配符的*含义不同,正则里要用.*表示任意字符序列:
#!/usr/bin/env python import xml.etree.ElementTree as ET, sys, os, re param = sys.argv[1] client = param.split('_')[0] market = param.split('_')[1] suffix = param.split('_')[2] # 编译正则:匹配符合基础规则,且不包含uat或test的文件名 regex_pattern = re.compile(rf"{re.escape(market)}.*2apex.*{re.escape(client)}.*\.xml$") # 额外的排除条件:文件名中不包含uat和test files_dir = '/some/dir' config_files = os.listdir(files_dir) for f in config_files: if regex_pattern.match(f) and 'uat' not in f and 'test' not in f: print(f)
或者也可以把排除逻辑写到正则里,用负向预查:
regex_pattern = re.compile(rf"{re.escape(market)}.*2apex.*{re.escape(client)}(?!.*uat)(?!.*test).*\.xml$")
然后循环里只需要判断regex_pattern.match(f)即可。
为什么你之前的尝试失败?
(?!uat):这是正则的负向预查语法,但fnmatch不支持正则的高级语法,只支持简单的shell通配符(*、?、[]),所以这个写法完全不被识别。re.compile(...)后直接赋值给toapex_pattern,然后传给fnmatch.fnmatch():fnmatch需要的是字符串格式的通配符,不是正则对象,所以会抛出类型错误。[^uat]:shell通配符里的[]是匹配单个字符,[^uat]表示匹配一个不是u、a、t的字符,而不是排除整个"uat"子串,所以无法实现你想要的排除效果。
内容的提问来源于stack exchange,提问作者kamokoba
相关产品推荐
相关产品推荐

