正则表达式多结果匹配问题:按起止字符分割字符串失效
format_..._end Segments Hey there! Let's get that regex working for you so you can pull out those three segments correctly.
Why Your Original Regex Isn't Cutting It
Your current regex (?:^|\s)format_(.*?)_end(?:\s|$) is looking for whitespace or string boundaries to bookend each segment—but your target string has no spaces. Each end directly runs into the next format, which breaks the regex's expectations:
- The first
endisn't followed by whitespace or the string end, so the regex can't finish matching the first segment. - Every subsequent
formatdoesn't have whitespace before it, so the regex completely ignores those instances.
The Correct Regex
Instead of relying on whitespace, we just need to match every chunk that starts with format_ and ends at the nearest _end. Here's the simple, effective regex you need:
format_.*?_end
Breakdown of How It Works
format_: Matches the literal start of each target segment..*?: Non-greedy match of any characters (this ensures we stop at the first_endinstead of skipping ahead to the last one)._end: Matches the literal end of each segment.
Example Code (Python)
Let's test this with your input string:
import re input_str = "format_abc_endformat_def_endformat_ghi_end" extracted_segments = re.findall(r'format_.*?_end', input_str) print(extracted_segments) # Output: ['format_abc_end', 'format_def_end', 'format_ghi_end']
Bonus: More Precise Match (If Your Use Case Allows)
If you know the middle part (between format_ and _end) will never contain underscores, you can use a stricter regex to avoid accidental matches:
format_[^_]*_end
This uses [^_]* to match only characters that aren't underscores, which adds an extra layer of safety for specific scenarios.
内容的提问来源于stack exchange,提问作者Erwin Vorenhout

