Python正则表达式问题:如何提取两个图片文件名
修正正则实现多文件名提取
问题场景
给定字符串:
line = '![[screenshotone.png]] and the next in the same line as ![[screenshottwo.jpg]]'
需要提取出screenshotone.png和screenshottwo.jpg组成列表,但使用以下代码时得到了错误结果:
import re output = re.findall('\[\[(.*)\]\]', line, re.I) print(output) # 输出: ['screenshotone.png]] and the next in the same line as ![[screenshottwo.jpg']
问题原因
原正则里的(.*)是贪婪匹配,会从第一个[[开始,尽可能多地匹配字符,直到找到最后一个]]才停止,所以把中间的无关文本也包含了进去。
两种修正方案
方案1:改用非贪婪匹配
把(.*)改成(.*?),让正则引擎尽可能少地匹配字符,遇到第一个]]就停止当前匹配:
import re line = '![[screenshotone.png]] and the next in the same line as ![[screenshottwo.jpg]]' output = re.findall(r'\[\[(.*?)\]\]', line, re.I) print(output) # 输出: ['screenshotone.png', 'screenshottwo.jpg']
方案2:使用排除式匹配
用[^\]]+替代.*,表示只匹配不是]的字符,这样不会跨过]]去匹配后面的内容:
import re line = '![[screenshotone.png]] and the next in the same line as ![[screenshottwo.jpg]]' output = re.findall(r'\[\[([^\]]+)\]\]', line, re.I) print(output) # 输出: ['screenshotone.png', 'screenshottwo.jpg']
说明
两种方案都能实现预期效果:
- 非贪婪匹配
.*?适合大多数场景,语法简洁; - 排除式匹配
[^\]]+更精准,避免意外匹配到特殊字符,稳定性更强。
内容的提问来源于stack exchange,提问作者nichas
相关产品推荐
相关产品推荐

