如何使用正则表达式正确提取x.set(对应括号内的完整内容(支持多行与嵌套括号)
Hey, let's fix that regex issue you're having! The problem with your current pattern is that it doesn't account for nested parentheses or properly balance opening/closing brackets—so it's either stopping too early or pulling in extra content like code outside the target parentheses.
The Correct Regex Solution
To reliably capture the content inside x.set(...) (including nested parentheses and cross-line content), we need a regex that can balance parentheses. Python's re module supports recursive patterns for this exact scenario:
import re # Your input string (with line breaks for clarity) value = """ver = '1.0' if x.set('1.2'): p = x.set('python_version', None) x = x.set('test_template', DEFAULT, p(x,b), z())""" # Regex pattern to match balanced parentheses inside x.set() pattern = r'(?<![0-9a-zA-Z_])x\.set\(\s*((?:[^()]|(?R))*?)\s*\)' # Extract the content inside each x.set() call raw_results = re.findall(pattern, value) # If you want to split each result into a list (like your expected output) # Note: This split works for simple cases—avoid if your args have commas inside strings! processed_results = [res.strip().split(', ') for res in raw_results] print("Raw captured content:") print(raw_results) # Output: ["'1.2'", "'python_version', None", "'test_template', DEFAULT, p(x,b), z()"] print("\nProcessed into lists:") print(processed_results) # Output: [["'1.2'"], ["'python_version'", "None"], ["'test_template'", "DEFAULT", "p(x,b)", "z()"]]
Breaking Down the Regex
Let's break down why this pattern works:
(?<![0-9a-zA-Z_]): Negative lookbehind to ensurex.setisn't part of a longer identifier (e.g., avoids matchingyx.set(...)).x\.set\(: Explicitly matchesx.set((we escape the.because it's a wildcard in regex).\s*: Matches any leading whitespace (including newlines) inside the parentheses.((?:[^()]|(?R))*?): The core of the pattern:(?:...): Non-capturing group for grouping logic without adding extra capture groups.[^()]: Matches any character that isn't an opening or closing parenthesis.|(?R): Recursively matches the entire regex pattern—this handles nested parentheses (likep(x,b)).*?: Non-greedy match to stop at the first matching closing parenthesis.
\s*\): Matches any trailing whitespace before the closing parenthesis.
A More Reliable Alternative for Python Code
If you're parsing valid Python code, using Python's built-in ast (Abstract Syntax Tree) module is far more robust than regex. It can handle edge cases like commas inside strings, complex nested expressions, and more:
import ast value = """ver = '1.0' if x.set('1.2'): p = x.set('python_version', None) x = x.set('test_template', DEFAULT, p(x,b), z())""" # Parse the code into an AST tree tree = ast.parse(value) results = [] # Traverse all nodes to find x.set() calls for node in ast.walk(tree): if (isinstance(node, ast.Call) and isinstance(node.func, ast.Attribute) and node.func.attr == 'set' and isinstance(node.func.value, ast.Name) and node.func.value.id == 'x'): # Convert each argument back to a string representation args = [ast.unparse(arg) for arg in node.args] results.append(args) print("AST-extracted results:") print(results) # Output: [["'1.2'"], ["'python_version'", 'None'], ["'test_template'", 'DEFAULT', 'p(x, b)', 'z()']]
This approach avoids all the pitfalls of regex when dealing with programming language syntax.
内容的提问来源于stack exchange,提问作者weis_ss

