正则技术需求:匹配最少层级括号内内容并排除内部嵌套括号
Hey there! Let's break down what's going wrong with your regex and how to fix this problem.
First, let's clarify your goal: you want to extract the content that's at the shallowest nesting level in a string of parentheses. That means:
- In
(((1)))(2),2is wrapped in only 1 layer of parentheses (shallowest), so we return2. - In
(((1)2)3),3is only wrapped by the outermost parentheses (level 1), while1is 3 levels deep—so we return3. - In
(1(2(3))),1is at level 1, which is shallower than2(level 2) and3(level 3)—so we return1.
Why Your Current Regex Isn't Working
Your regex \(((?>[^\(\)]+|(?R))*)\) is a recursive pattern that matches complete nested parenthesis structures (like the entire (((1))) or (2) block in your first example). It doesn't track the nesting level of individual content fragments—it just grabs whole blocks of nested parentheses instead of targeting the shallowest inner content. That's why it's not aligning with your expected results.
Solution 1: Track Nesting Level with Code
Regex isn't the most reliable tool for dynamic nesting level tracking. A simple script that manually tracks parenthesis depth will work far better. Here's a Python example:
def get_shallowest_content(s): current_depth = 0 content_segments = [] current_segment = [] for char in s: if char == '(': # Save any ongoing content before increasing depth if current_segment: content_segments.append((''.join(current_segment), current_depth)) current_segment = [] current_depth += 1 elif char == ')': # Save any ongoing content before decreasing depth if current_segment: content_segments.append((''.join(current_segment), current_depth)) current_segment = [] current_depth -= 1 else: # Add non-parenthesis characters to the current segment current_segment.append(char) # Save the last segment if there's any left if current_segment: content_segments.append((''.join(current_segment), current_depth)) if not content_segments: return "" # Find the shallowest depth (smallest number) shallowest_depth = min(depth for _, depth in content_segments) # Return the first segment at the shallowest depth for seg, depth in content_segments: if depth == shallowest_depth: return seg.strip() return ""
Test Results:
get_shallowest_content("(((1)))(2)")→ returns"2"get_shallowest_content("(((1)2)3)")→ returns"3"get_shallowest_content("(1(2(3)))")→ returns"1"
Solution 2: Regex Approach (Limited Use Cases)
If you specifically need a regex solution, it works best for simpler scenarios. Here are two targeted patterns:
Match non-nested parentheses (content with no inner parentheses):
\(([^()]+)\)This will directly match
(2)in your first example, extracting"2".Match outermost non-nested content:
For cases like(((1)2)3)or(1(2(3))), use this pattern to grab content in the outermost parentheses that isn't wrapped in any inner brackets:(?<=\()([^()]+)(?=[^()]*\))This will extract
"3"from(((1)2)3)and"1"from(1(2(3))).
Note: This regex method won't handle mixed cases perfectly (e.g., if there's both non-nested parentheses and outermost content), so the code-based approach is more robust.
内容的提问来源于stack exchange,提问作者Kevin A.S.

