Ply Lex令牌逻辑解析及自定义令牌规则技术咨询
Great questions! Let's break this down step by step—Ply's lexer behavior and regex fundamentals are key to building exactly the tokenizer you need.
Ply's lexer works by first compiling all your token regex patterns into state machines. When processing input text, it scans from the current position, trying to find the longest possible match against any of your defined token patterns. Priority here is two-fold: longer matches always win, and if two patterns match the same length, the one defined earlier in your code takes precedence.
For your t_NUMBER_LITERAL example:
- The regex
r'\d+'targets one or more consecutive digits. - When the lexer finds a match, it runs the associated function, converts the matched string value to an integer with
t.value = int(t.value), and returns the token so the parser can use it.
Let's walk through each of your requested patterns with working, practical code examples:
1. Matching Variables
Variables typically start with a letter or underscore, followed by letters, numbers, or underscores. Here's how to define that token:
def t_VARIABLE(t): r'[a-zA-Z_][a-zA-Z0-9_]*' # Optional: Add a lookup if you need to handle reserved keywords (like 'if', 'else') # t.type = reserved_words.get(t.value, 'VARIABLE') return t
The regex [a-zA-Z_] matches the valid starting character, and [a-zA-Z0-9_]* matches zero or more subsequent valid characters for the variable name.
2. Matching Strings Wrapped in Double Quotes
To match any string enclosed in double quotes (including empty strings), use this rule. We'll capture the content inside the quotes and strip the surrounding quotes from the token value:
def t_STRING_LITERAL(t): r'"([^"]*)"' # Extract the content inside the quotes (group 1 of the regex) t.value = t.group(1) return t
If you need to handle escaped quotes inside the string (like "He said \"Hello!\""), modify the regex to account for escape sequences:
def t_STRING_LITERAL(t): r'"(\\.|[^"])*"' t.value = t.group(1).replace('\\"', '"') # Unescape the quoted characters return t
The \\. matches any escaped character (like \" or \\), and [^"]* matches any non-quote character outside escape sequences.
3. Ignoring Comments Inside Curly Braces
Ply has a special convention for ignored tokens: prefix the rule name with t_ignore_. For non-nested {} comments, define this:
def t_ignore_COMMENT(t): r'\{[^}]*\}' # No return statement needed — Ply automatically skips this token
Note: This regex only works for non-nested curly braces. If your comments can have nested {} (like { Outer { inner } comment }), you'll need a stateful lexer since basic regex can't handle nested structures. Here's a quick implementation:
# Define a new exclusive state for comments states = ( ('comment', 'exclusive'), ) # Enter comment state when we see '{' def t_comment_start(t): r'\{' t.lexer.push_state('comment') t.lexer.comment_depth = 1 # Track nesting depth # Handle nested braces inside comments def t_comment_curly(t): r'[\{\}]' if t.value == '{': t.lexer.comment_depth += 1 else: t.lexer.comment_depth -= 1 # Exit comment state when depth returns to 0 if t.lexer.comment_depth == 0: t.lexer.pop_state() # Ignore all other characters inside comments def t_comment_ignore(t): r'[^{}]+' pass # Error handling for unclosed comments def t_comment_error(t): print(f"Unclosed comment at line {t.lineno}") t.lexer.skip(1)
r'\d+' Only Match Digits? In Python's regex engine (which Ply uses), \d is a shorthand character class that strictly matches decimal digits (0-9). The + quantifier means "match one or more of the preceding element". Combined, r'\d+' will match one or more consecutive digit characters, and nothing else. If you wanted to match Unicode digits (like Arabic or Devanagari numerals), you'd need a broader character class, but for standard use cases, \d is explicitly for 0-9 digits.
内容的提问来源于stack exchange,提问作者Berecz Balázs

