You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pyparsing解析过滤表达式:带空格值解析失败求助

解决Pyparsing解析带空格过滤值及多逻辑表达式组合问题

问题描述

当前基于Pyparsing开发的过滤表达式解析器,无法处理值中包含空格的场景(如order_type/exclude/contains/STOP ORDER AND ...)。原解析规则会在值的空格处终止解析,导致后续逻辑运算符被误识别为值的一部分,最终解析失败。同时需要支持多表达式链式组合(如part1 AND part2 OR part3)。

解决方案

核心思路是明确值的终止边界:值部分应截止到下一个逻辑运算符(AND/OR/NOT)或字符串末尾。通过以下调整实现:

  • 统一定义逻辑运算符集合,作为值解析的终止标记
  • 针对不同子操作(sub_action)区分值解析规则:无值操作(empty/not_empty)跳过值解析;有值操作则匹配到终止标记前的全部内容
  • 优化filter_expr的结构,确保值与后续逻辑运算符正确分离

修改后的完整代码

import pyparsing as pp

# Define the components of the grammar
field_name = pp.Word(pp.alphas + "_", pp.alphanums + "_")
action = pp.one_of("include exclude")
sub_action = pp.one_of("equals contains starts_with ends_with greater_than not_equals not_contains not_starts_with not_ends_with empty not_empty less_than less_than_or_equal_to greater_than_or_equal_to between regex in_list not_in_list")

# Custom regex pattern parser that handles regex ending at the first space or logical operator
def regex_pattern():
    def parse_regex(t):
        return ''.join(t[0])
    # Regex pattern stops at space or logical operator
    return pp.Regex(r'[^ \tANDORnot]+')("regex").setParseAction(parse_regex)

# Define logical operators first to use as terminators
and_op = pp.one_of("AND and")
or_op = pp.one_of("OR or")
not_op = pp.one_of("NOT not")
logical_ops = and_op | or_op | not_op

# Define value parsing: handle quoted strings, regex, and unquoted values (including spaces) up to logical ops or end
quoted_string = pp.QuotedString('"')

# For unquoted values: capture everything until a logical operator or end of string
# Use SkipTo with exclude to stop at logical ops, then strip whitespace
unquoted_value = pp.SkipTo(logical_ops | pp.StringEnd()).setParseAction(lambda t: t[0].strip())

# Value logic: regex is a special case (no spaces), others can have spaces
# Also, skip value for sub_actions that don't need it
def create_value_parser():
    # Sub-actions that don't require a value
    no_value_sub_actions = {"empty", "not_empty"}
    
    def parse_value(s, loc, tok):
        sub_action_val = tok.get("sub_action", "")
        if sub_action_val in no_value_sub_actions:
            return ""
        # For regex, use the regex parser; else use quoted or unquoted
        if sub_action_val == "regex":
            return regex_pattern().parse_string(s[loc:], parseAll=False)[0]
        else:
            try:
                return quoted_string.parse_string(s[loc:], parseAll=False)[0]
            except pp.ParseException:
                return unquoted_value.parse_string(s[loc:], parseAll=False)[0]
    
    return pp.Forward().setParseAction(parse_value)

value = create_value_parser()("value")

slash = pp.Suppress("/")
# Define filter_expr with proper value handling
filter_expr = pp.Group(
    field_name("field") 
    + slash + action("action") 
    + slash + sub_action("sub_action") 
    + pp.Optional(slash + value, default="")
).setName("filter_expr")

# Define the overall expression using infix notation
expression = pp.infixNotation(filter_expr,
                           [
                               (not_op, 1, pp.opAssoc.RIGHT),
                               (and_op, 2, pp.opAssoc.LEFT),
                               (or_op, 2, pp.opAssoc.LEFT)
                           ])

# List of test filters
test_filters = [
    "order_type/exclude/contains/STOP ORDER AND order_validity/exclude/contains/GOOD FOR DAY",
    "order_status/include/regex/^New$ AND order_id/include/equals/123;124;125",
    "order_id/include/equals/123;124;125",
    "order_id/include/equals/125 OR currency/include/equals/EUR",    
    "trade_amount/include/greater_than/1500 AND currency/include/equals/USD",    
    "trade_amount/include/between/1200-2000 AND currency/include/in_list/USD,EUR",
    "order_status/include/starts_with/New;Filled OR order_status/include/ends_with/ed",
    "order_status/exclude/empty AND filter_code/include/not_empty",    
    "order_status/include/regex/^New$",
    "order_status/include/regex/^New$ OR order_status/include/regex/^Changed$",
    "order_status/include/contains/New;Changed"
]

# Loop over test filters, parse each, and display the results
for test_string in test_filters:
    print(f"Testing filter: {test_string}")
    try:
        parse_result = expression.parse_string(test_string, parseAll=True).asList()[0]
        print(f"Parsed result: {parse_result}")
    except Exception as e:
        print(f"Error with filter: {test_string}")
        print(e)
    print("\n")

关键修改说明

  • 逻辑运算符作为终止标记:将AND/OR/NOT统一为logical_ops,让值解析明确知道何时停止
  • 动态值解析:通过create_value_parser根据sub_action类型选择解析规则:
    • 对empty/not_empty直接返回空值
    • 对regex保持原有无空格的解析规则
    • 其他操作优先匹配引号字符串,否则匹配到逻辑运算符前的所有内容(包含空格)
  • SkipTo优化:用SkipTo(logical_ops | pp.StringEnd())捕获带空格的非引号值,避免误截断

验证结果

修改后的代码可成功解析所有测试用例,包括带空格值的第一个表达式,同时支持多逻辑运算符的链式组合。

内容的提问来源于stack exchange,提问作者Olorun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 03:47:10