.NET单行正则需求:移除首字符及最后下划线后内容(兼容特殊字符)
Let's break down why your original regex is failing and fix it right away.
The Problem with Your Current Regex
Your pattern (?<=a)[\w ]+(?=_) only matches word characters (\w = letters, digits, underscores) and spaces. When it hits special characters like -, ~, or =, it can't include them in the match—so it stops early and grabs the last valid segment of \w/space characters before the final underscore. That's why you're getting rmenia instead of the full desired string in your second test case.
The Fix
We need a regex that captures everything after the leading a and before the final underscore, regardless of what special characters are in between. Here are two reliable options for single-line mode:
Option 1: Lookarounds (No Capture Groups)
(?<=a).*(?=_[^_]*$)
(?<=a): Positive lookbehind to ensure we start right after the leadinga..*: Greedily matches any character (single-line mode lets this work across line breaks if needed, though your use case is single-line).(?=_[^_]*$): Positive lookahead to ensure we stop at the last underscore (it checks that after the underscore, there are no more underscores until the end of the string).
Option 2: Capture Group (More Intuitive for Some Engines)
If lookarounds feel tricky, this version uses a capture group to directly extract the desired content:
^a(.*)_[^_]*$
^a: Matches the leadingaat the start of the string.(.*): Captures everything between theaand the final underscore._[^_]*$: Matches the final underscore plus all characters after it (which we ignore, keeping only the captured group).
Testing the Fix
Let's run both regexes against your test cases:
- For
aPersonal Protective Equipment_REV2.docx:- Both patterns return
Personal Protective Equipment(correct).
- Both patterns return
- For
aFreight Forwarder Standard Operating Procedure - Armenia_REV1.docx:- Both patterns return
Freight Forwarder Standard Operating Procedure - Armenia(fixed the original error).
- Both patterns return
Key Takeaway
The original regex's character class [\w ] was too restrictive. By switching to .* (which matches any character) and ensuring we stop at the last underscore (not just any underscore), we handle all special characters correctly while meeting your requirements.
内容的提问来源于stack exchange,提问作者24K Gold

