You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyParsing如何区分sect关键字与普通参数 同时解析全局配置和sect块

问题描述

需要解析如下格式的testp.txt文件:

title                               = Test Suite A;
timeout                             = 10000
exp_delay                           = 500;
log                                 = TRUE;

sect
{
    type                            = typeA;
    name                            = "HelloWorld";
    output_log                      = "c:\test\out.log";
};

sect
{
    name = "GoodbyeAll";
    type = typeB;
    comm1_req = 0xDEADBEEF;
    comm1_resp = (int, 1234366);
};

文件结构为:开头是全局参数区域,之后跟随任意数量的sect块。目前单独解析参数或单独解析sect块均可正常工作,但同时包含两种结构的文件解析失败。
原始代码中param规则会误匹配sect关键字,添加~Literal("sect")排除后仍报错:

Exception raised:Found unwanted token, "sect", found '\n'  (at char 188), (line:4, col:56)
Expected end of text, found 's'  (at char 190), (line:6, col:1)

已尝试将final规则改为final = ZeroOrMore(param) + ZeroOrMore(node)解决了部分问题,需要更结构化的实现方式直接输出符合要求的字典,包含全局参数和所有sect条目。

解决方案

核心问题说明

  1. 原始final规则使用|选择符,仅支持全参数或全sect块的单一结构文件,两种结构共存时会匹配失败。
  2. 原始node规则定义漏写右括号,存在语法错误。

优化后代码

from pyparsing import *
from pathlib import Path

# 基础规则定义
command_req = Word(alphanums)
command_resp = Group("(" + delimitedList(Word(alphanums + "x")) + ")").setParseAction(lambda t: tuple(t[0][1:-1]))
kW = Word(alphas+'_', alphanums+'_') | command_req | command_resp
keyName = ~Literal("sect") + Word(alphas+'_', alphanums+'_') + FollowedBy("=")
# 优化keyValue解析,避免单值返回列表
keyValue = (
    dblQuotedString.setParseAction(removeQuotes) 
    | Combine(OneOrMore(kW, stopOn=LineEnd() | ";"))
)
param = dictOf(keyName, Suppress("=") + keyValue + Optional(Suppress(";")))
# 修复node规则的括号缺失,给sect内容加结构化解析
node = Group(
    Suppress("sect") + Suppress("{") 
    + Dict(OneOrMore(param)) 
    + Suppress("};")
)
# 定义最终规则,设置结果名直接输出结构化字典
final = Dict(
    ZeroOrMore(param).setResultsName("global_params")
    + ZeroOrMore(node).setResultsName("sects")
)

p = Path(__file__).with_name("testp.txt")
with open(p) as f:
    try:
        x = final.parseFile(f, parseAll=True)
        dx = x.asDict()
        print(dx)
    except ParseException as pe:
        print(pe)

输出结果示例

解析后得到的字典结构如下:

{
    "global_params": {
        "title": "Test Suite A",
        "timeout": "10000",
        "exp_delay": "500",
        "log": "TRUE"
    },
    "sects": [
        {
            "type": "typeA",
            "name": "HelloWorld",
            "output_log": "c:\\test\\out.log"
        },
        {
            "name": "GoodbyeAll",
            "type": "typeB",
            "comm1_req": "0xDEADBEEF",
            "comm1_resp": ("int", "1234366")
        }
    ]
}

内容的提问来源于stack exchange,提问作者SimpleOne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 19:27:03