为何PyParsing解析器抛出非预期ParseException而非自定义FilterException?
问题描述
我正在实现一款解析器,用于处理输入字符串、提取组件、验证并生成SQLAlchemy查询,目前在过滤器验证环节遇到以下问题:
- 已自定义
FilterException,预期当过滤器为"w"时触发该异常 - 实际运行时,解析器抛出的是
ParseException: Expected end of text, found 'and'这类非预期异常,而非自定义的FilterException
过滤器定义如下:
filter_term = Combine(Optional(space) + Word(alphas) + Optional(space)).set_results_name("filter").set_parse_action( filter_validator).set_name("filter")
计划为过滤器添加额外验证,通过含别名的字典限定可用过滤器词,示例字典:
{ "animal": "animal", "dog": "animal", "cat": "animal", "pet": "animal" }
当前的验证逻辑仅做简单检查,但未生效:
if t[0] == "w": raise FilterException("Invalid filter")
完整代码
from pyparsing import Word, Combine, Optional, DelimitedList, alphanums, Suppress, Group, one_of, alphas, \ CaselessLiteral, infix_notation, opAssoc, OneOrMore, Keyword, CaselessKeyword, pyparsing_common, Forward, \ ParseException, ParseSyntaxException, ZeroOrMore class OrOperation: def __init__(self, instring, loc, toks): raise ParseException(instring, loc, "invalid OR given") class AndOperation: def __init__(self, instring, loc, toks): raise ParseException(instring, loc, "invalid AND given") class FilterException(ParseException): def __init__(self, pstr): super().__init__(pstr) def filter_validator(s, l, t): if t[0] == "w": raise FilterException("Invalid filter") # utils: comma = Suppress(",") space = Suppress(" ") lbrace = Suppress("(") rbrace = Suppress(")") and_operator = Suppress(CaselessKeyword("AND")) or_operator = CaselessKeyword("OR") search_parser = Forward().set_name("search_expression") literal_value = Forward().set_name("literal_value").set_results_name("literal_value") delimited_list_delim = Optional(comma + Optional(space)) delimited_list = DelimitedList(literal_value, delim=delimited_list_delim).set_parse_action( lambda tokens: ", ".join(tokens)) string_literal = Word(alphanums + "_") wildcard_literal = Combine(string_literal + "*").set_parse_action(lambda tokens: tokens[0].replace("*", "?")) delimited_list_literal = lbrace + delimited_list + rbrace filter_term = Combine(Optional(space) + Word(alphas) + Optional(space)).set_results_name("filter").set_parse_action( filter_validator).set_name("filter") literal_value <<= delimited_list_literal | wildcard_literal | string_literal equals_operator = one_of("= :") comparison_operator = one_of("> >= < <= ") not_equals_operator = CaselessLiteral("!=") contains_operator = CaselessLiteral("~").set_parse_action(lambda tokens: "LIKE") not_contains_operator = CaselessLiteral("!~").set_parse_action(lambda tokens: "NOT LIKE") operator = equals_operator | not_equals_operator | contains_operator | not_contains_operator | comparison_operator operator_term = Combine(Optional(space) + operator + Optional(space)).set_results_name("operator") expression_term = Group(filter_term + operator_term + literal_value).set_parse_action(filter_validator) | Group( literal_value) search_parser <<= infix_notation(expression_term, [ (and_operator, 2, opAssoc.LEFT, lambda instring, loc, toks: AndOperation(instring, loc, toks)), (or_operator, 2, opAssoc.LEFT, lambda instring, loc, toks: OrOperation(instring, loc, toks)) ]) try: result = search_parser.parse_string("w~(a, b c, d)") print(result.dump()) except FilterException as e: print("Filter failed:", e) search_parser.run_tests(''' asas was* (as, b,c d) ((as, b,c d)) w=a w=a* w=(a, b c, d) w:(a, b c, d) w!=(a, b c, d) w~(a, b c, d) w!~(a, b c, d) w>=(a, b c, d) a>=(a, b c, d) and a=(a, b c, d) w>=(a, b c, d) and w=(a, b c, d) and w=(a, b c, d) w>=(a, b c, d) or (w=(a, b c, d) and w=(a, b c, d)) (w>=(a, b c, d) or w!~(a, b c, d)) or (w=(a, b c, d) and w=(a, b c, d)) w>=(a, b c, d) or w!~(a, b c, d) or (w=(a, b c, d) and w=(a, b c, d)) w>=(a, b c, d) or w!~(a, b c, d) or w=(a, b c, d) and w=(a, b c, d) a>=(a, b c, d) and w!~(a, b c, d) or w=(a, b c, d) and w!=(a, b c, d) ''')
问题分析与解决方案
核心问题
expression_term重复绑定验证函数:你给Group(filter_term + operator_term + literal_value)也绑定了filter_validator,但这里的t是整个分组的结果,并非单个过滤器字符串,导致验证逻辑完全偏离预期。FilterException初始化参数错误:ParseException的构造函数需要传入(instring, loc, message)三个参数,你仅传入单个字符串,导致异常位置信息错误,干扰了pyparsing的异常处理流程。- 解析回溯导致异常偏移:当
filter_term的验证抛出异常后,解析器会尝试回溯匹配expression_term的另一个分支(literal_value),后续遇到"and"等符号时,就会抛出与过滤器验证无关的语法异常。
修复步骤
1. 修正FilterException构造函数
class FilterException(ParseException): def __init__(self, instring, loc, message): super().__init__(instring, loc, message)
2. 移除expression_term上多余的验证绑定
filter_validator仅需绑定在filter_term上即可:
expression_term = Group(filter_term + operator_term + literal_value) | Group(literal_value)
3. 调整验证函数逻辑
利用set_results_name设置的"filter"字段获取值,同时处理空格问题:
def filter_validator(s, l, t): filter_val = t.filter.strip() if filter_val == "w": raise FilterException(s, l, "Invalid filter")
4. 实现字典式过滤器别名验证(可选)
按照计划添加别名验证逻辑:
ALLOWED_FILTERS = { "animal": "animal", "dog": "animal", "cat": "animal", "pet": "animal" } def filter_validator(s, l, t): filter_val = t.filter.strip().lower() if filter_val not in ALLOWED_FILTERS: raise FilterException(s, l, f"Invalid filter: {filter_val}") # 将别名替换为实际字段名 t["filter"] = ALLOWED_FILTERS[filter_val]
5. 测试验证
修改后,输入"w~(a, b c, d)"会直接抛出FilterException;对于包含"and"的测试用例,会在第一个非法过滤器位置就触发异常,而非等到解析到"and"时才报错。
内容的提问来源于stack exchange,提问作者PyJaGa
相关产品推荐
相关产品推荐

