正则表达式替换transpiler令牌时如何排除引号内文本?
Hey there! I get that you're building a transpiler that converts custom language tokens to Python ones using regex substitutions, but you're stuck on skipping text inside quotes—totally relatable, since regex can be tricky with context like that. Let me walk you through a couple of practical approaches that solve this problem.
Approach 1: Use Positive Lookahead to Exclude Quotes
The core idea here is to add a regex assertion that ensures your target token isn't inside a pair of quotes. We can use a positive lookahead to verify that any quotes appearing after the token are properly closed (meaning the token is outside all quoted sections).
Example Code
import re # Your replacement rules: (pattern, translated_token) reps = [ (r'VAR(?=(?:[^"]|"[^"]*")*$)', 'def'), (r'PRINT(?=(?:[^"]|"[^"]*")*$)', 'print') ] # Sample input from your custom language transpiled = """ VAR my_func(): PRINT "Hello, VAR!" PRINT some_var """ # Apply substitutions for pattern, translated in reps: transpiled = re.sub(pattern, translated, transpiled, flags=re.UNICODE) print(transpiled)
Output
def my_func(): print "Hello, VAR!" print some_var
How the Regex Works
The (?=(?:[^"]|"[^"]*")*$) part is the key:
(?:[^"]|"[^"]*")matches either a non-quote character, or a full quoted string (from"to").- The
*$ensures this pattern repeats until the end of the line, meaning all quotes after the token are paired and closed. So the token can't be inside an unclosed quote.
Approach 2: Capture Quotes and Skip Them (More Flexible)
If you have multiple tokens to replace or need to handle edge cases better, using a custom replacement function is more robust. This approach captures quoted text separately and leaves it untouched, while replacing tokens outside quotes.
Example Code
import re def replace_token_outside_quotes(match): # If we matched a quoted string, return it as-is if match.group(1): return match.group(1) # Otherwise, map the custom token to Python's equivalent token = match.group(2) token_map = { 'VAR': 'def', 'PRINT': 'print' } return token_map.get(token, token) # Fallback to original if no match transpiled = """ VAR my_func(): PRINT "Hello, VAR!" PRINT some_var """ # Regex pattern: match either quoted text OR target tokens pattern = r'(".*?")|(VAR|PRINT)' transpiled = re.sub(pattern, replace_token_outside_quotes, transpiled, flags=re.UNICODE) print(transpiled)
Handling Edge Cases
Both methods assume you're using unescaped double quotes. If your language supports escaped quotes (like "He said \"Hello\""), you'll need to adjust the regex to account for escaped characters. For example, replace "[^"]*" with "[^"\\]*(?:\\.[^"\\]*)*"—this matches quoted strings that include escaped quotes.
内容的提问来源于stack exchange,提问作者Lucca

