You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何PyParsing解析器抛出非预期ParseException而非自定义FilterException?

问题描述

我正在实现一款解析器,用于处理输入字符串、提取组件、验证并生成SQLAlchemy查询,目前在过滤器验证环节遇到以下问题:

  • 已自定义FilterException,预期当过滤器为"w"时触发该异常
  • 实际运行时,解析器抛出的是ParseException: Expected end of text, found 'and'这类非预期异常,而非自定义的FilterException

过滤器定义如下:

filter_term = Combine(Optional(space) + Word(alphas) + Optional(space)).set_results_name("filter").set_parse_action(
    filter_validator).set_name("filter")

计划为过滤器添加额外验证,通过含别名的字典限定可用过滤器词,示例字典:

{
    "animal": "animal",
    "dog": "animal",
    "cat": "animal",
    "pet": "animal"
}

当前的验证逻辑仅做简单检查,但未生效:

if t[0] == "w":
    raise FilterException("Invalid filter")
完整代码
from pyparsing import Word, Combine, Optional, DelimitedList, alphanums, Suppress, Group, one_of, alphas, \
    CaselessLiteral, infix_notation, opAssoc, OneOrMore, Keyword, CaselessKeyword, pyparsing_common, Forward, \
    ParseException, ParseSyntaxException, ZeroOrMore


class OrOperation:
    def __init__(self, instring, loc, toks):
        raise ParseException(instring, loc, "invalid OR given")


class AndOperation:
    def __init__(self, instring, loc, toks):
        raise ParseException(instring, loc, "invalid AND given")


class FilterException(ParseException):
    def __init__(self, pstr):
        super().__init__(pstr)


def filter_validator(s, l, t):
    if t[0] == "w":
        raise FilterException("Invalid filter")


# utils:
comma = Suppress(",")
space = Suppress(" ")
lbrace = Suppress("(")
rbrace = Suppress(")")
and_operator = Suppress(CaselessKeyword("AND"))
or_operator = CaselessKeyword("OR")

search_parser = Forward().set_name("search_expression")
literal_value = Forward().set_name("literal_value").set_results_name("literal_value")

delimited_list_delim = Optional(comma + Optional(space))
delimited_list = DelimitedList(literal_value, delim=delimited_list_delim).set_parse_action(
    lambda tokens: ", ".join(tokens))

string_literal = Word(alphanums + "_")
wildcard_literal = Combine(string_literal + "*").set_parse_action(lambda tokens: tokens[0].replace("*", "?"))
delimited_list_literal = lbrace + delimited_list + rbrace

filter_term = Combine(Optional(space) + Word(alphas) + Optional(space)).set_results_name("filter").set_parse_action(
    filter_validator).set_name("filter")
literal_value <<= delimited_list_literal | wildcard_literal | string_literal

equals_operator = one_of("= :")
comparison_operator = one_of("> >= < <= ")
not_equals_operator = CaselessLiteral("!=")
contains_operator = CaselessLiteral("~").set_parse_action(lambda tokens: "LIKE")
not_contains_operator = CaselessLiteral("!~").set_parse_action(lambda tokens: "NOT LIKE")
operator = equals_operator | not_equals_operator | contains_operator | not_contains_operator | comparison_operator
operator_term = Combine(Optional(space) + operator + Optional(space)).set_results_name("operator")
expression_term = Group(filter_term + operator_term + literal_value).set_parse_action(filter_validator) | Group(
    literal_value)

search_parser <<= infix_notation(expression_term,
                                 [
                                     (and_operator, 2, opAssoc.LEFT,
                                      lambda instring, loc, toks: AndOperation(instring, loc, toks)),
                                     (or_operator, 2, opAssoc.LEFT,
                                      lambda instring, loc, toks: OrOperation(instring, loc, toks))
                                 ])

try:
    result = search_parser.parse_string("w~(a, b c, d)")
    print(result.dump())
except FilterException as e:
    print("Filter failed:", e)

search_parser.run_tests('''
asas
was*
(as, b,c d)
((as, b,c d))
w=a
w=a*
w=(a, b c, d)
w:(a, b c, d)
w!=(a, b c, d)
w~(a, b c, d)
w!~(a, b c, d)
w>=(a, b c, d)
a>=(a, b c, d) and a=(a, b c, d)
w>=(a, b c, d) and w=(a, b c, d) and w=(a, b c, d)
w>=(a, b c, d) or (w=(a, b c, d) and w=(a, b c, d))
(w>=(a, b c, d) or w!~(a, b c, d))  or (w=(a, b c, d) and w=(a, b c, d))
w>=(a, b c, d) or w!~(a, b c, d)  or (w=(a, b c, d) and w=(a, b c, d))
w>=(a, b c, d) or w!~(a, b c, d)  or w=(a, b c, d) and w=(a, b c, d)
a>=(a, b c, d) and w!~(a, b c, d)  or w=(a, b c, d) and w!=(a, b c, d)
''')
问题分析与解决方案

核心问题

  1. expression_term重复绑定验证函数:你给Group(filter_term + operator_term + literal_value)也绑定了filter_validator,但这里的t是整个分组的结果,并非单个过滤器字符串,导致验证逻辑完全偏离预期。
  2. FilterException初始化参数错误:ParseException的构造函数需要传入(instring, loc, message)三个参数,你仅传入单个字符串,导致异常位置信息错误,干扰了pyparsing的异常处理流程。
  3. 解析回溯导致异常偏移:当filter_term的验证抛出异常后,解析器会尝试回溯匹配expression_term的另一个分支(literal_value),后续遇到"and"等符号时,就会抛出与过滤器验证无关的语法异常。

修复步骤

1. 修正FilterException构造函数

class FilterException(ParseException):
    def __init__(self, instring, loc, message):
        super().__init__(instring, loc, message)

2. 移除expression_term上多余的验证绑定

filter_validator仅需绑定在filter_term上即可:

expression_term = Group(filter_term + operator_term + literal_value) | Group(literal_value)

3. 调整验证函数逻辑

利用set_results_name设置的"filter"字段获取值,同时处理空格问题:

def filter_validator(s, l, t):
    filter_val = t.filter.strip()
    if filter_val == "w":
        raise FilterException(s, l, "Invalid filter")

4. 实现字典式过滤器别名验证(可选)

按照计划添加别名验证逻辑:

ALLOWED_FILTERS = {
    "animal": "animal",
    "dog": "animal",
    "cat": "animal",
    "pet": "animal"
}

def filter_validator(s, l, t):
    filter_val = t.filter.strip().lower()
    if filter_val not in ALLOWED_FILTERS:
        raise FilterException(s, l, f"Invalid filter: {filter_val}")
    # 将别名替换为实际字段名
    t["filter"] = ALLOWED_FILTERS[filter_val]

5. 测试验证

修改后,输入"w~(a, b c, d)"会直接抛出FilterException;对于包含"and"的测试用例,会在第一个非法过滤器位置就触发异常,而非等到解析到"and"时才报错。


内容的提问来源于stack exchange,提问作者PyJaGa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 18:54:51