You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从右侧构建非贪婪正则表达式?提取指定文件名的目标内容

解决正则捕获最后一个下划线后、.pdf前内容的问题

Got it, let's break down what's going on here and fix your regex!

Your current pattern _.*?\.pdf starts matching at the first underscore in the string, and the non-greedy .*? stops at the first .pdf it finds—so it ends up grabbing everything from that first underscore to the end of the filename, which isn't what you want. We need to target the last underscore specifically, then capture the content between it and .pdf.

Here are a few reliable approaches to get your desired 12a3 fragment:

方法1:匹配最后一个下划线后的非下划线字符

Use a negated character class [^_]+ to match one or more characters that aren't underscores. This ensures we only grab content after the final underscore (since there are no underscores left after that point):

import re
s = 'ab9c_xy8z_12a3.pdf'
m = re.search(r'_([^_]+)\.pdf', s)
print(m.group(1))  # Output: '12a3'
  • _: Matches the final underscore
  • ([^_]+): Captures one or more characters that aren't underscores (this is your target fragment)
  • \.pdf: Matches the file extension (the backslash escapes the dot, which is a special regex character)

方法2:贪婪匹配到最后一个下划线

Use .*_ to greedily match everything up to the last underscore, then capture the content before .pdf:

m = re.search(r'.*_([^.]+)\.pdf', s)
print(m.group(1))  # Output: '12a3'
  • .*_: The greedy .* will match as much as possible, so it stops at the last underscore in the string
  • ([^.]+): Captures one or more characters that aren't dots (perfect for avoiding the .pdf extension)
  • \.pdf: Matches the file extension

方法3:零宽断言(无需捕获组)

If you prefer to avoid using capture groups, you can use positive lookbehind and lookahead assertions to directly match the target fragment:

m = re.search(r'(?<=_)[^.]+(?=\.pdf)', s)
print(m.group())  # Output: '12a3'
  • (?<=_): Positive lookbehind—ensures the character before our match is an underscore
  • [^.]+: Matches one or more non-dot characters
  • (?=\.pdf): Positive lookahead—ensures the characters after our match are .pdf

All three methods will give you exactly the 12a3 fragment you're looking for. Pick the one that makes the most sense for your use case!

内容的提问来源于stack exchange,提问作者harshal1618

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:41:21