如何为res = 3 + x_sum*11这类简单表达式构建分词器?
Let's walk through building a tokenizer for your simple expression use case step by step. I'll use Python for examples since it's widely used and has great regex support, but the core logic translates to other languages too.
Step 1: Clarify Token Rules
First, let's restate the token definitions to make sure we cover everything clearly:
- Integer Literals: One or more digits (0-9), e.g.,
3,11 - Identifiers: Start with a letter (a-z, A-Z) or underscore, followed by letters, digits, or underscores, e.g.,
res,x_sum - Operators: Specific symbols like
=,+,*(we can easily extend this later) - Whitespace: Skip leading/trailing whitespace entirely, and ignore whitespace between tokens.
Step 2: Simple Implementation with Regex finditer
Regex is perfect here because it lets us define patterns for each token type and scan the string efficiently. Here's a straightforward implementation:
import re def tokenize(expression): # Remove leading/trailing whitespace first cleaned_expr = expression.strip() # Define regex patterns grouped by token type token_regex = re.compile( r'([a-zA-Z_][a-zA-Z0-9_]*)|(\d+)|([=+*])' ) tokens = [] for match in token_regex.finditer(cleaned_expr): # Check which group matched to determine token type if match.group(1): tokens.append(('IDENTIFIER', match.group(1))) elif match.group(2): # Convert to integer if you need the numeric value, or keep as string tokens.append(('INTEGER', int(match.group(2)))) elif match.group(3): tokens.append(('OPERATOR', match.group(3))) return tokens # Test with your example xpr = "res = 3 + x_sum*11" print(tokenize(xpr)) # Output: [('IDENTIFIER', 'res'), ('OPERATOR', '='), ('INTEGER', 3), ('OPERATOR', '+'), ('IDENTIFIER', 'x_sum'), ('OPERATOR', '*'), ('INTEGER', 11)]
How This Works:
strip()handles leading/trailing whitespace as required.- The regex uses three capture groups: one for identifiers, one for integers, one for operators.
finditer()scans the string and returns matches in order. We check which group has a value to assign the correct token type.- Middle whitespace is automatically skipped because it doesn't match any of our token patterns.
Step 3: Robust Implementation with re.Scanner
For better error handling (like catching invalid characters) and cleaner whitespace handling, use re.Scanner—it's designed exactly for tokenization tasks:
import re def tokenize_with_validation(expression): cleaned_expr = expression.strip() # Define scanner rules: (pattern, handler function) scanner = re.Scanner([ # Skip whitespace entirely (r'\s+', lambda scanner, token: None), # Match identifiers (r'[a-zA-Z_][a-zA-Z0-9_]*', lambda scanner, token: ('IDENTIFIER', token)), # Match integers and convert to numeric type (r'\d+', lambda scanner, token: ('INTEGER', int(token))), # Match supported operators (r'[=+*]', lambda scanner, token: ('OPERATOR', token)), ]) tokens, remaining = scanner.scan(cleaned_expr) # If there's any unprocessed text, it's an invalid character if remaining: raise ValueError(f"Unrecognized character(s) in expression: '{remaining}'") return tokens # Test with valid input print(tokenize_with_validation("res = 3 + x_sum*11")) # Test with invalid input (will throw an error) # tokenize_with_validation("res = 3 + x_sum@11")
Key Improvements:
- Explicitly skips whitespace using a handler that returns
None. - Validates the entire expression—if any characters don't match a token pattern (like
@), it raises a clear error. - More modular: adding new operators is as simple as updating the operator regex (e.g.,
[+\-*/=]to add subtraction and division).
Step 4: Extending the Tokenizer
Want to add more operators (like -, /) or support for floating-point numbers? Just adjust the regex patterns:
- For more operators: Update the operator pattern to
[+\-*/=](note:-goes at the start/end of the character class to avoid being interpreted as a range). - For floats: Add a new pattern like
\d+\.\d+and a corresponding group/handler.
内容的提问来源于stack exchange,提问作者user3639447

