You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pyparsing解析多段多行归档文件描述块的解决方案

问题

刚接触pyparsing,已实现单个归档文件描述块的解析,但无法处理包含开头可忽略的file_metadata块、多个归档项及项间换行的真实数据,最终目标是解析所有归档项存入数据库。

已实现的单块解析代码:

import pyparsing as pp

one_archive = \
"""archive (
    name "something wicked this way comes.zip"
    file ( name wicked.exe size 140084 date 2022/12/24 23:32:00 crc B2CF5E58 )
    file ( name readme.txt size 1704 date 2022/12/24 23:32:00 crc 37F73AEE )
)
"""
pp.ParserElement.set_default_whitespace_chars(' \t')

EOL = pp.LineEnd().suppress()
start_of_archive_block = pp.LineStart() + pp.Keyword('archive (') + EOL
end_of_archive_block = pp.LineStart() + ')' + EOL

archive_filename = pp.LineStart() \
    + pp.Keyword('name').suppress() \
    + pp.Literal('"').suppress() \
    + pp.SkipTo(pp.Literal('"')).set_results_name("archive_name") \
    + pp.Literal('"').suppress() \
    + EOL

field_elem = pp.Keyword('name').suppress() + pp.SkipTo(pp.Literal(' size')).set_results_name("filename") \
    ^ pp.Keyword('size').suppress() + pp.SkipTo(pp.Literal(' date')).set_results_name("size") \
    ^ pp.Keyword('date').suppress() + pp.SkipTo(pp.Literal(' crc')).set_results_name("date") \
    ^ pp.Keyword('crc').suppress() + pp.SkipTo(pp.Literal(' )')).set_results_name("crc")
fields = field_elem * 4

filerow = pp.LineStart() \
    + pp.Literal('file (').suppress() \
    + fields \
    + pp.Literal(')').suppress() \
    + EOL

archive = start_of_archive_block.suppress() \
    + archive_filename \
    + pp.OneOrMore(pp.Group(filerow)) \
    + end_of_archive_block.suppress()

archive.parse_string(one_archive, parse_all=True)

需要处理的真实数据示例:

realistic_data = \
"""
file_metadata (
    description: blah blah etc.
    author: john doe
    version: 0.99
)

archive (
    name "something wicked this way comes.zip"
    file ( name wicked.exe size 140084 date 2022/12/24 23:32:00 crc B2CF5E58 )
    file ( name readme.txt size 1704 date 2022/12/24 23:32:00 crc 37F73AEE )
)

archive (
    name "naughty or nice.zip"
    file ( name naughty.exe size 187232 date 2021/8/4 10:19:55 crc 638BC6AA )
    file ( name nice.exe size 298234 date 2021/8/4 10:19:56 crc 99FD31AE )
    file ( name whatever.jpg size 25603 date 2021/8/5 11:03:09 crc ABFAC314 )
)
"""

需求:跳过file_metadata块,解析多个archive项及项间换行。

解决方案

1. 定义可忽略的file_metadata解析器

匹配整个file_metadata块并标记为可忽略,避免干扰后续解析:

# 匹配file_metadata块内部所有内容,直到闭合的),并跳过块后的换行
file_metadata = pp.Keyword("file_metadata (").suppress() \
    + pp.SkipTo(pp.LineStart() + ")").suppress() \
    + pp.LineStart() + ")".suppress() \
    + pp.ZeroOrMore(pp.LineEnd())

2. 优化归档项解析逻辑

原代码依赖固定后续关键词的SkipTo容易出错,改成更健壮的字段匹配规则:

  • 文件名:匹配除)外的所有可打印字符
  • 大小:匹配纯数字串
  • 日期时间:组合日期和时间部分
  • CRC:匹配十六进制字符串

调整后的单个文件项解析器:

filename = pp.Keyword("name").suppress() + pp.Word(pp.printables, exclude_chars=")").set_results_name("filename")
size = pp.Keyword("size").suppress() + pp.Word(pp.nums).set_results_name("size")
datetime_str = pp.Combine(pp.Word(pp.nums + "/") + pp.Word(pp.nums + ":"))
date = pp.Keyword("date").suppress() + datetime_str.set_results_name("date")
crc = pp.Keyword("crc").suppress() + pp.Word(pp.hexnums.upper()).set_results_name("crc")

filerow = pp.LineStart() \
    + pp.Keyword("file (").suppress() \
    + filename + size + date + crc \
    + pp.Literal(")").suppress() \
    + pp.LineEnd().suppress()

3. 组合完整解析器,支持多个归档块

允许开头可选的file_metadata块,后续匹配一个或多个归档块,自动忽略块间空白:

# 单个归档块解析器
archive = pp.LineStart() + pp.Keyword("archive (").suppress() + pp.LineEnd().suppress() \
    + pp.Keyword("name").suppress() + pp.QuotedString('"').set_results_name("archive_name") + pp.LineEnd().suppress() \
    + pp.OneOrMore(pp.Group(filerow)).set_results_name("files") \
    + pp.LineStart() + ")".suppress() \
    + pp.ZeroOrMore(pp.LineEnd())

# 完整解析规则:可选元数据块 + 多个归档块 + 字符串结束
full_parser = pp.Optional(file_metadata) + pp.OneOrMore(pp.Group(archive)) + pp.StringEnd()

完整修改后的代码

import pyparsing as pp

# 配置默认空白字符(空格和制表符)
pp.ParserElement.set_default_whitespace_chars(' \t')

# 1. 定义可忽略的file_metadata块解析器
file_metadata = pp.Keyword("file_metadata (").suppress() \
    + pp.SkipTo(pp.LineStart() + ")").suppress() \
    + pp.LineStart() + ")".suppress() \
    + pp.ZeroOrMore(pp.LineEnd())

# 2. 定义单个file项的解析规则
filename = pp.Keyword("name").suppress() + pp.Word(pp.printables, exclude_chars=")").set_results_name("filename")
size = pp.Keyword("size").suppress() + pp.Word(pp.nums).set_results_name("size")
datetime_str = pp.Combine(pp.Word(pp.nums + "/") + pp.Word(pp.nums + ":"))
date = pp.Keyword("date").suppress() + datetime_str.set_results_name("date")
crc = pp.Keyword("crc").suppress() + pp.Word(pp.hexnums.upper()).set_results_name("crc")

filerow = pp.LineStart() \
    + pp.Keyword("file (").suppress() \
    + filename + size + date + crc \
    + pp.Literal(")").suppress() \
    + pp.LineEnd().suppress()

# 3. 定义单个archive块解析器
archive = pp.LineStart() + pp.Keyword("archive (").suppress() + pp.LineEnd().suppress() \
    + pp.Keyword("name").suppress() + pp.QuotedString('"').set_results_name("archive_name") + pp.LineEnd().suppress() \
    + pp.OneOrMore(pp.Group(filerow)).set_results_name("files") \
    + pp.LineStart() + ")".suppress() \
    + pp.ZeroOrMore(pp.LineEnd())

# 4. 完整解析器
full_parser = pp.Optional(file_metadata) + pp.OneOrMore(pp.Group(archive)) + pp.StringEnd()

# 测试真实数据
realistic_data = \
"""
file_metadata (
    description: blah blah etc.
    author: john doe
    version: 0.99
)

archive (
    name "something wicked this way comes.zip"
    file ( name wicked.exe size 140084 date 2022/12/24 23:32:00 crc B2CF5E58 )
    file ( name readme.txt size 1704 date 2022/12/24 23:32:00 crc 37F73AEE )
)

archive (
    name "naughty or nice.zip"
    file ( name naughty.exe size 187232 date 2021/8/4 10:19:55 crc 638BC6AA )
    file ( name nice.exe size 298234 date 2021/8/4 10:19:56 crc 99FD31AE )
    file ( name whatever.jpg size 25603 date 2021/8/5 11:03:09 crc ABFAC314 )
)
"""

# 解析并处理结果
results = full_parser.parse_string(realistic_data)

# 遍历结果,可直接对接数据库插入逻辑
for archive_block in results:
    archive_name = archive_block.archive_name
    print(f"归档文件: {archive_name}")
    for file in archive_block.files:
        print(f"  文件: {file.filename}, 大小: {file.size}, 日期: {file.date}, CRC: {file.crc}")
    # 此处添加数据库操作代码,比如:
    # cursor.execute("INSERT INTO archives (name) VALUES (?)", (archive_name,))
    # for file in archive_block.files:
    #     cursor.execute("INSERT INTO files (archive_id, filename, size, date, crc) VALUES (?, ?, ?, ?, ?)", 
    #                   (archive_id, file.filename, file.size, file.date, file.crc))

关键说明

  • pp.Optional(file_metadata)自动跳过开头的元数据块,不影响后续解析
  • pp.OneOrMore(pp.Group(archive))支持多个归档块,Group()将每个归档的结果封装为独立组,方便遍历处理
  • 优化后的字段匹配规则不再依赖固定后续关键词,提升了解析的健壮性
  • ZeroOrMore(pp.LineEnd())处理块间的换行和空白,确保解析不受格式换行影响

内容的提问来源于stack exchange,提问作者JK Laiho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 20:24:50