如何通过正则表达式提取测试字符串中的编码与对应文本?
Got it, let's figure out why your current regex isn't working and fix it up! The problem with your existing pattern is that the (?<title>[\s\S]+) part is greedy—it will match absolutely everything from the first code's closing bracket all the way to the end of the string, which is why you can't get separate entries for each code-title pair.
The Solution Regex
Here's a pattern that will capture exactly the results you want:
\((?<code>\d{3})\)\s?(?<title>.*?)(?=\s*\(\d{3}\)|$)
Breakdown of the Pattern
Let's walk through each part so you understand how it works:
\(: Matches the opening round bracket (we can adjust this if you need to support square brackets too—more on that later)(?<code>\d{3}): Captures exactly 3 digits as thecodegroup (your original pattern allowed 0-3 digits, but your expected results all use 3-digit codes, so this is more precise)\): Matches the closing round bracket\s?: Matches an optional single space after the closing bracket(?<title>.*?): Non-greedily captures the title content. The?makes it stop matching as soon as it hits the next boundary, instead of gobbling everything up.(?=\s*\(\d{3}\)|$): A positive lookahead that tells the regex to stop matching the title when it sees either:- Optional whitespace followed by another
(XXX)code pattern, or - The end of the string (
$)
- Optional whitespace followed by another
Testing Against Your Sample String
When you run this regex on your test string:(003) Recoverable salaries and allowances (General) (700) General non-recurrent, (001) lamslkdf; (999) ajkndsfk
You'll get exactly the expected matches:
code: 003,title: Recoverable salaries and allowances (General)code: 700,title: General non-recurrent,code: 001,title: lamslkdf;code: 999,title: ajkndsfk
If You Need to Support Square Brackets
If your actual data might have square brackets like [003] instead of just (003), adjust the pattern to include both bracket types:
[\(\[]?(?<code>\d{3})[\)\]]\s?(?<title>.*?)(?=\s*[\(\[]\d{3}[\)\]]|$)
Why Your Original Regex Failed
To recap, two key issues:
- Greedy matching:
[\s\S]+matches every character (including newlines) until the end of the string—so your first match would capture003as code, and the entire rest of the string as title. - Loose code matching:
[0-9]{0,3}allows 0, 1, 2, or 3 digits, which could lead to unexpected captures if there are shorter numeric sequences in your text. Using\d{3}ensures you only get 3-digit codes.
内容的提问来源于stack exchange,提问作者Ken Hui

