You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ply Lex令牌逻辑解析及自定义令牌规则技术咨询

Great questions! Let's break this down step by step—Ply's lexer behavior and regex fundamentals are key to building exactly the tokenizer you need.

Ply Lex Tokenizer Underlying Logic

Ply's lexer works by first compiling all your token regex patterns into state machines. When processing input text, it scans from the current position, trying to find the longest possible match against any of your defined token patterns. Priority here is two-fold: longer matches always win, and if two patterns match the same length, the one defined earlier in your code takes precedence.

For your t_NUMBER_LITERAL example:

  • The regex r'\d+' targets one or more consecutive digits.
  • When the lexer finds a match, it runs the associated function, converts the matched string value to an integer with t.value = int(t.value), and returns the token so the parser can use it.
Custom Token Rules for Your Use Cases

Let's walk through each of your requested patterns with working, practical code examples:

1. Matching Variables

Variables typically start with a letter or underscore, followed by letters, numbers, or underscores. Here's how to define that token:

def t_VARIABLE(t):
    r'[a-zA-Z_][a-zA-Z0-9_]*'
    # Optional: Add a lookup if you need to handle reserved keywords (like 'if', 'else')
    # t.type = reserved_words.get(t.value, 'VARIABLE')
    return t

The regex [a-zA-Z_] matches the valid starting character, and [a-zA-Z0-9_]* matches zero or more subsequent valid characters for the variable name.

2. Matching Strings Wrapped in Double Quotes

To match any string enclosed in double quotes (including empty strings), use this rule. We'll capture the content inside the quotes and strip the surrounding quotes from the token value:

def t_STRING_LITERAL(t):
    r'"([^"]*)"'
    # Extract the content inside the quotes (group 1 of the regex)
    t.value = t.group(1)
    return t

If you need to handle escaped quotes inside the string (like "He said \"Hello!\""), modify the regex to account for escape sequences:

def t_STRING_LITERAL(t):
    r'"(\\.|[^"])*"'
    t.value = t.group(1).replace('\\"', '"')  # Unescape the quoted characters
    return t

The \\. matches any escaped character (like \" or \\), and [^"]* matches any non-quote character outside escape sequences.

3. Ignoring Comments Inside Curly Braces

Ply has a special convention for ignored tokens: prefix the rule name with t_ignore_. For non-nested {} comments, define this:

def t_ignore_COMMENT(t):
    r'\{[^}]*\}'
    # No return statement needed — Ply automatically skips this token

Note: This regex only works for non-nested curly braces. If your comments can have nested {} (like { Outer { inner } comment }), you'll need a stateful lexer since basic regex can't handle nested structures. Here's a quick implementation:

# Define a new exclusive state for comments
states = (
    ('comment', 'exclusive'),
)

# Enter comment state when we see '{'
def t_comment_start(t):
    r'\{'
    t.lexer.push_state('comment')
    t.lexer.comment_depth = 1  # Track nesting depth

# Handle nested braces inside comments
def t_comment_curly(t):
    r'[\{\}]'
    if t.value == '{':
        t.lexer.comment_depth += 1
    else:
        t.lexer.comment_depth -= 1
        # Exit comment state when depth returns to 0
        if t.lexer.comment_depth == 0:
            t.lexer.pop_state()

# Ignore all other characters inside comments
def t_comment_ignore(t):
    r'[^{}]+'
    pass

# Error handling for unclosed comments
def t_comment_error(t):
    print(f"Unclosed comment at line {t.lineno}")
    t.lexer.skip(1)
Why Does r'\d+' Only Match Digits?

In Python's regex engine (which Ply uses), \d is a shorthand character class that strictly matches decimal digits (0-9). The + quantifier means "match one or more of the preceding element". Combined, r'\d+' will match one or more consecutive digit characters, and nothing else. If you wanted to match Unicode digits (like Arabic or Devanagari numerals), you'd need a broader character class, but for standard use cases, \d is explicitly for 0-9 digits.

内容的提问来源于stack exchange,提问作者Berecz Balázs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:50:04