Python正则匹配含多可选组不同长度复杂文件名的问题
正则修改方案
问题原因
原来的正则存在以下错误:
- 可选的tile、provider分组没有和前置分隔下划线绑定,
_?(?P<tile>\w+)?这种写法导致贪婪模式下\w+会吞噬后续provider字段甚至日期前的内容,最终两个分组都匹配不到值 - 字段匹配规则不够严谨,
\w只能匹配字母、数字、下划线,无法匹配其他特殊字符 - 后缀前的
.作为正则元字符没有转义,存在匹配任意字符的风险
修正后的正则与代码
import re pattern = r'^(?P<product>\w{3})' \ r'_(?P<country>\w{2})' \ r'-(?P<region>[^_]+)' \ r'(?:_(?P<tile>[^_]+))?' \ r'(?:_(?P<provider>[^_]+))?' \ r'_(?P<year>\d{4})' \ r'(?P<month>\d{2})' \ r'(?P<day>\d{2})' \ r'_(?P<hour>\d{2})' \ r'(?P<minute>\d{2})' \ r'(?P<second>\d{2})' \ r'_(?P<sensor>\w{5})' \ r'_(?P<res_unit>km|m|cm)' \ r'(?P<resolution>\d{3,4})' \ r'(?:_(?P<bittype>\d{1,2}bit))?' \ r'\.(?P<format>\w+)$' p = re.compile(pattern) # 测试文件名 fn1 = 'FOO_is-atest_123456_COMPANY_20190729_153343_SATEL_m0001_32bit.tif' fn2 = 'FOO_is-atest_COMPANY_20190729_153343_SATEL_m0001_32bit.tif' fn3 = 'FOO_is-atest_COMPANY_20190729_153343_SATEL_m0001.tif' fn4 = 'FOO_is-atest_32tnt_20211125_120005_SATEL_m0001.tif' fn5 = 'FOO_is-atest_20211125_120005_SATEL_cm070.tif' fn6 = 'FOO_is-atest_20211125_120005_SATEL_cm070_32bit.tif' print(p.match(fn1).group('tile'), p.match(fn1).group('provider')) print(p.match(fn2).group('provider'), p.match(fn2).group('bittype')) print(p.match(fn3).group('provider'), p.match(fn3).group('resolution')) print(p.match(fn4).group('tile'), p.match(fn4).group('year')) print(p.match(fn5).group('provider'), p.match(fn5).group('resolution')) print(p.match(fn6).group('provider'), p.match(fn6).group('bittype'))
输出结果
123456 COMPANY COMPANY 32bit COMPANY 0001 32tnt 2021 None 070 None 32bit
说明
- 用
(?:_xxx)?的非捕获分组结构,把分隔下划线和字段绑定,只有存在下划线加对应内容时才会匹配分组,避免吞噬问题 - 用
[^_]+匹配字段内容,匹配到下一个下划线为止,支持除下划线外的所有字符 - 增加了
^和$锚定整个文件名的首尾,避免多余字符干扰匹配 - 转义了后缀前的
.,严格匹配点号
内容的提问来源于stack exchange,提问作者s6hebern
相关产品推荐
相关产品推荐

