You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用ANTLR4解析Velocity变量?

Hey there! Let's work through splitting plain text and Velocity variables with a lexer. First, let's restate the Velocity variable syntax clearly since that's our anchor:

变量的简写形式以“$”开头,后跟VTL标识符。VTL标识符必须以字母(a..z或A..Z)开头,其余字符只能是字母(a..z、A..Z)、数字(0..9)或下划线(“_”)。

Core Lexer Strategy

The key here is to prioritize matching valid Velocity variables first, then treat everything else as plain text. This avoids edge cases where a $ isn't part of a variable (like $123 or $@).

Regex for Velocity Variables

We can use this regex to match valid variables:

\$[a-zA-Z][a-zA-Z0-9_]*

Breakdown:

  • \$: Escaped literal $ (since $ is a special character in regex)
  • [a-zA-Z]: Mandatory starting character (must be upper/lowercase letter)
  • [a-zA-Z0-9_]*: Optional trailing characters (letters, numbers, underscores)

Example Lexer Implementation (Python PLY)

If you're using Python's PLY library, here's a working skeleton that builds on your existing code:

import ply.lex as lex

# Define token types
tokens = (
    'VELOCITY_VAR',
    'PLAIN_TEXT'
)

# Match valid Velocity variables (priority over plain text)
def t_VELOCITY_VAR(t):
    r'\$[a-zA-Z][a-zA-Z0-9_]*'
    return t

# Match plain text: either non-$ characters, or $ followed by non-letter (invalid variable)
def t_PLAIN_TEXT(t):
    r'[^$]+|\$(?![a-zA-Z])'
    return t

# Handle newlines (adjust based on your needs)
def t_newline(t):
    r'\n+'
    t.lexer.lineno += len(t.value)

# Error handling for invalid characters
def t_error(t):
    print(f"Illegal character: '{t.value[0]}'")
    t.lexer.skip(1)

# Initialize lexer
lexer = lex.lex()

# Test with sample input
test_content = "Hi $first_name! Your discount code is $save_2024. Note: $123 is NOT a variable."
lexer.input(test_content)

# Print token output
for token in lexer:
    print(f"Token Type: {token.type} | Value: '{token.value}'")

Key Edge Cases Handled

  • $123 or $@ are treated as plain text (since they don't start with a letter after $)
  • Variables with underscores and numbers (like $save_2024) are correctly identified
  • Plain text containing $ (that isn't part of a variable) is preserved

Test Output

Running the sample code will produce this output:

Token Type: PLAIN_TEXT | Value: 'Hi '
Token Type: VELOCITY_VAR | Value: '$first_name'
Token Type: PLAIN_TEXT | Value: '! Your discount code is '
Token Type: VELOCITY_VAR | Value: '$save_2024'
Token Type: PLAIN_TEXT | Value: '. Note: '
Token Type: PLAIN_TEXT | Value: '$123'
Token Type: PLAIN_TEXT | Value: ' is NOT a variable.'

内容的提问来源于stack exchange,提问作者WAKU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:09:50