如何用正则表达式提取含嵌套方括号字符串中的目标内容?
Got it, let's break down why your original regex isn't working and fix it. Your pattern \[.*?\] uses a non-greedy match, which stops at the first closing bracket it finds—this is the inner ] from [0], hence the partial results like [Test.A[0.
Here are a few solid solutions to extract exactly Test.A[0], Test.B[0], etc.:
Solution 1: Directly Match the Target Pattern
If you know your target strings follow the format Test.[Letter][Number], you can skip dealing with the outer brackets entirely and match the exact pattern:
Test\.\w+\[\d+\]
Test\.: Matches the literal "Test." (the backslash escapes the dot, which otherwise acts as a wildcard)\w+: Matches one or more word characters (covers letters like A/B/C/D)\[\d+\]: Matches[followed by one or more digits, then]
This regex will directly find all occurrences of your desired strings without needing to parse the outer brackets.
Solution 2: Extract Content from Outer Brackets
If you want to pull whatever is inside the outer square brackets (regardless of the prefix, as long as there's no ' inside), use this regex with a capture group:
\[([^']+)\]
\[: Matches the opening outer bracket([^']+): Capture group that matches any character except'(since your outer brackets are always followed by'in the input)\]: Matches the closing outer bracket
The capture group will give you exactly the content inside each outer bracket pair (e.g., Test.A[0] from '[Test.A[0]]').
Solution 3: Exact Match for Your Specific Format
For a regex tightly tailored to your input structure:
\[(Test\.[A-Z]\[\d\])\]
\[(...)\]: Captures everything inside the outer bracketsTest\.: Matches "Test."[A-Z]: Matches a single uppercase letter (A-D in your example)\[\d\]: Matches[followed by a single digit, then]
Capture group 1 will return your desired strings perfectly.
Example Usage (Python)
import re input_str = "('[Test.A[0]]' <>'' OR '[Test.B[0]]' <>'' OR '[Test.C[0]]' <>'' OR '[Test.D[0]]' <> '')" # Using Solution 1 matches = re.findall(r'Test\.\w+\[\d+\]', input_str) print(matches) # Output: ['Test.A[0]', 'Test.B[0]', 'Test.C[0]', 'Test.D[0]'] # Using Solution 2 matches = re.findall(r'\[([^']+)\]', input_str) print(matches) # Output: ['Test.A[0]', 'Test.B[0]', 'Test.C[0]', 'Test.D[0]']
内容的提问来源于stack exchange,提问作者Edward

