You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Lexer生成测试数据测试Parser的相关技术疑问

关于转译器Parser测试的疑问

我正在编写一个小型转译器,已为Lexer编写了大量测试,但编写Parser测试难度更高,因为Parser需要接收List[Token],而Lexer仅接收str。

不使用Lexer生成List[Token]的测试示例

def some_test_name():
    first_decimal_token: Token = Token(DECIMAL_TOKEN_TYPE, "86")
    multiply_token: Token = Token(MULTIPLY_TOKEN_TYPE, "*")
    second_decimal_token: Token = Token(DECIMAL_TOKEN_TYPE, "3")

    tokens: List[Token] = [
        first_decimal_token, multiply_token, second_decimal_token,
        end_of_file_token
    ]

    first_factor_node: FactorNode = FactorNode(first_decimal_token.value)
    multiply_operator: ArithmeticOperator = ArithmeticOperator.MULTIPLY
    second_factor_node: FactorNode = FactorNode(second_decimal_token.value)

    term_node: TermNode = TermNode(first_factor_node, multiply_operator,
                                   second_factor_node)

    expected_output_tokens: List[Token] = [end_of_file_token]
    expected_output: NodeSuccess = NodeSuccess(expected_output_tokens, term_node)

    node_output: NodeResult = parse_tokens_for_term(tokens)

    assert isinstance(node_output, NodeSuccess)
    assert expected_output == node_output

使用Lexer的测试示例

def using_da_lexer():
    INPUT: str = "86 * 3"
    
    lexer_result: List[Token] | LexerFailure = lexer(INPUT)
    assert isinstance(lexer_result, list)

    first_decimal_token: Token = lexer_result[0]
    second_decimal_token: Token = lexer_result[2]

    first_factor_node: FactorNode = FactorNode(first_decimal_token.value)
    multiply_operator: ArithmeticOperator = ArithmeticOperator.MULTIPLY
    second_factor_node: FactorNode = FactorNode(second_decimal_token.value)

    term_node: TermNode = TermNode(first_factor_node, multiply_operator,
                                   second_factor_node)

    expected_output_tokens: List[Token] = [end_of_file_token]
    expected_output: NodeSuccess = NodeSuccess(expected_output_tokens, term_node)

    node_output: NodeSuccess | NodeFailure = parse_tokens_for_term(tokens)

    assert isinstance(node_output, NodeSuccess)
    assert expected_output == node_output

1. 在Parser测试中使用Lexer是否存在问题?

没问题,但要区分场景:

  • 如果Lexer已经经过全面测试,用它生成Token列表是安全的,还能避免手动构造Token时的类型错误、拼写错误。
  • 但不要用Lexer测试Parser的边界异常场景——比如测试Parser对非法Token序列(缺失运算符、Token类型不匹配等)的处理,此时必须手动构造Token列表,因为Lexer不会生成不符合语法的Token序列,无法覆盖这类测试点。

2. 哪种测试方法更具可扩展性?测试复杂表达式时哪种更省心?

使用Lexer的方法扩展性更强,测试复杂表达式时也更省心:

  • 测试"86 * 2 + 10 / 3"这类复杂表达式时,手动构造每个Token会非常繁琐,重复代码多且易出错。用Lexer只需要写输入字符串,Token生成交给已验证过的Lexer即可,代码量大幅减少。
  • 新增表达式类型(比如带括号、变量)时,字符串输入的方式能快速生成测试用例,而手动构造Token需要逐个调整每个Token的类型和值,效率极低。
  • 手动构造Token的方法适合单元级小测试,比如测试单个Term或Factor的解析逻辑,能精准控制输入的Token序列,聚焦Parser的特定分支。

3. 有没有更优的Parser测试方案,还是编写全面测试本就需要大量代码?

有几个优化方案可以减少重复代码:

  • 封装Token构造工具函数:比如写make_tokens(types_and_values)函数,传入类型和值的列表,自动生成包含EOF的Token序列,减少手动创建每个Token的代码。
  • 封装预期AST构造工具函数:比如写make_term_node(left_factor, op, right_factor)这类函数,快速生成预期的AST节点,不用每次手动实例化所有子节点。
  • 参数化测试:利用测试框架的参数化功能(比如Python的pytest.mark.parametrize),把输入字符串和预期AST(或预期结果)做成参数对,一次运行多个测试用例,避免重复编写测试函数。
  • 快照测试:如果AST可以序列化为稳定格式(比如JSON),可以用快照测试工具,自动保存第一次运行的AST结果,后续测试直接对比快照,不用手动构造预期节点。

当然,编写全面的Parser测试本来就需要不少代码,但通过工具函数和参数化测试能大幅减少重复工作,提升效率。

内容的提问来源于stack exchange,提问作者8SIXSector

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 19:13:12