使用pyparsing解析多段多行归档文件描述块的解决方案
问题
刚接触pyparsing,已实现单个归档文件描述块的解析,但无法处理包含开头可忽略的file_metadata块、多个归档项及项间换行的真实数据,最终目标是解析所有归档项存入数据库。
已实现的单块解析代码:
import pyparsing as pp one_archive = \ """archive ( name "something wicked this way comes.zip" file ( name wicked.exe size 140084 date 2022/12/24 23:32:00 crc B2CF5E58 ) file ( name readme.txt size 1704 date 2022/12/24 23:32:00 crc 37F73AEE ) ) """ pp.ParserElement.set_default_whitespace_chars(' \t') EOL = pp.LineEnd().suppress() start_of_archive_block = pp.LineStart() + pp.Keyword('archive (') + EOL end_of_archive_block = pp.LineStart() + ')' + EOL archive_filename = pp.LineStart() \ + pp.Keyword('name').suppress() \ + pp.Literal('"').suppress() \ + pp.SkipTo(pp.Literal('"')).set_results_name("archive_name") \ + pp.Literal('"').suppress() \ + EOL field_elem = pp.Keyword('name').suppress() + pp.SkipTo(pp.Literal(' size')).set_results_name("filename") \ ^ pp.Keyword('size').suppress() + pp.SkipTo(pp.Literal(' date')).set_results_name("size") \ ^ pp.Keyword('date').suppress() + pp.SkipTo(pp.Literal(' crc')).set_results_name("date") \ ^ pp.Keyword('crc').suppress() + pp.SkipTo(pp.Literal(' )')).set_results_name("crc") fields = field_elem * 4 filerow = pp.LineStart() \ + pp.Literal('file (').suppress() \ + fields \ + pp.Literal(')').suppress() \ + EOL archive = start_of_archive_block.suppress() \ + archive_filename \ + pp.OneOrMore(pp.Group(filerow)) \ + end_of_archive_block.suppress() archive.parse_string(one_archive, parse_all=True)
需要处理的真实数据示例:
realistic_data = \ """ file_metadata ( description: blah blah etc. author: john doe version: 0.99 ) archive ( name "something wicked this way comes.zip" file ( name wicked.exe size 140084 date 2022/12/24 23:32:00 crc B2CF5E58 ) file ( name readme.txt size 1704 date 2022/12/24 23:32:00 crc 37F73AEE ) ) archive ( name "naughty or nice.zip" file ( name naughty.exe size 187232 date 2021/8/4 10:19:55 crc 638BC6AA ) file ( name nice.exe size 298234 date 2021/8/4 10:19:56 crc 99FD31AE ) file ( name whatever.jpg size 25603 date 2021/8/5 11:03:09 crc ABFAC314 ) ) """
需求:跳过file_metadata块,解析多个archive项及项间换行。
解决方案
1. 定义可忽略的file_metadata解析器
匹配整个file_metadata块并标记为可忽略,避免干扰后续解析:
# 匹配file_metadata块内部所有内容,直到闭合的),并跳过块后的换行 file_metadata = pp.Keyword("file_metadata (").suppress() \ + pp.SkipTo(pp.LineStart() + ")").suppress() \ + pp.LineStart() + ")".suppress() \ + pp.ZeroOrMore(pp.LineEnd())
2. 优化归档项解析逻辑
原代码依赖固定后续关键词的SkipTo容易出错,改成更健壮的字段匹配规则:
- 文件名:匹配除
)外的所有可打印字符 - 大小:匹配纯数字串
- 日期时间:组合日期和时间部分
- CRC:匹配十六进制字符串
调整后的单个文件项解析器:
filename = pp.Keyword("name").suppress() + pp.Word(pp.printables, exclude_chars=")").set_results_name("filename") size = pp.Keyword("size").suppress() + pp.Word(pp.nums).set_results_name("size") datetime_str = pp.Combine(pp.Word(pp.nums + "/") + pp.Word(pp.nums + ":")) date = pp.Keyword("date").suppress() + datetime_str.set_results_name("date") crc = pp.Keyword("crc").suppress() + pp.Word(pp.hexnums.upper()).set_results_name("crc") filerow = pp.LineStart() \ + pp.Keyword("file (").suppress() \ + filename + size + date + crc \ + pp.Literal(")").suppress() \ + pp.LineEnd().suppress()
3. 组合完整解析器,支持多个归档块
允许开头可选的file_metadata块,后续匹配一个或多个归档块,自动忽略块间空白:
# 单个归档块解析器 archive = pp.LineStart() + pp.Keyword("archive (").suppress() + pp.LineEnd().suppress() \ + pp.Keyword("name").suppress() + pp.QuotedString('"').set_results_name("archive_name") + pp.LineEnd().suppress() \ + pp.OneOrMore(pp.Group(filerow)).set_results_name("files") \ + pp.LineStart() + ")".suppress() \ + pp.ZeroOrMore(pp.LineEnd()) # 完整解析规则:可选元数据块 + 多个归档块 + 字符串结束 full_parser = pp.Optional(file_metadata) + pp.OneOrMore(pp.Group(archive)) + pp.StringEnd()
完整修改后的代码
import pyparsing as pp # 配置默认空白字符(空格和制表符) pp.ParserElement.set_default_whitespace_chars(' \t') # 1. 定义可忽略的file_metadata块解析器 file_metadata = pp.Keyword("file_metadata (").suppress() \ + pp.SkipTo(pp.LineStart() + ")").suppress() \ + pp.LineStart() + ")".suppress() \ + pp.ZeroOrMore(pp.LineEnd()) # 2. 定义单个file项的解析规则 filename = pp.Keyword("name").suppress() + pp.Word(pp.printables, exclude_chars=")").set_results_name("filename") size = pp.Keyword("size").suppress() + pp.Word(pp.nums).set_results_name("size") datetime_str = pp.Combine(pp.Word(pp.nums + "/") + pp.Word(pp.nums + ":")) date = pp.Keyword("date").suppress() + datetime_str.set_results_name("date") crc = pp.Keyword("crc").suppress() + pp.Word(pp.hexnums.upper()).set_results_name("crc") filerow = pp.LineStart() \ + pp.Keyword("file (").suppress() \ + filename + size + date + crc \ + pp.Literal(")").suppress() \ + pp.LineEnd().suppress() # 3. 定义单个archive块解析器 archive = pp.LineStart() + pp.Keyword("archive (").suppress() + pp.LineEnd().suppress() \ + pp.Keyword("name").suppress() + pp.QuotedString('"').set_results_name("archive_name") + pp.LineEnd().suppress() \ + pp.OneOrMore(pp.Group(filerow)).set_results_name("files") \ + pp.LineStart() + ")".suppress() \ + pp.ZeroOrMore(pp.LineEnd()) # 4. 完整解析器 full_parser = pp.Optional(file_metadata) + pp.OneOrMore(pp.Group(archive)) + pp.StringEnd() # 测试真实数据 realistic_data = \ """ file_metadata ( description: blah blah etc. author: john doe version: 0.99 ) archive ( name "something wicked this way comes.zip" file ( name wicked.exe size 140084 date 2022/12/24 23:32:00 crc B2CF5E58 ) file ( name readme.txt size 1704 date 2022/12/24 23:32:00 crc 37F73AEE ) ) archive ( name "naughty or nice.zip" file ( name naughty.exe size 187232 date 2021/8/4 10:19:55 crc 638BC6AA ) file ( name nice.exe size 298234 date 2021/8/4 10:19:56 crc 99FD31AE ) file ( name whatever.jpg size 25603 date 2021/8/5 11:03:09 crc ABFAC314 ) ) """ # 解析并处理结果 results = full_parser.parse_string(realistic_data) # 遍历结果,可直接对接数据库插入逻辑 for archive_block in results: archive_name = archive_block.archive_name print(f"归档文件: {archive_name}") for file in archive_block.files: print(f" 文件: {file.filename}, 大小: {file.size}, 日期: {file.date}, CRC: {file.crc}") # 此处添加数据库操作代码,比如: # cursor.execute("INSERT INTO archives (name) VALUES (?)", (archive_name,)) # for file in archive_block.files: # cursor.execute("INSERT INTO files (archive_id, filename, size, date, crc) VALUES (?, ?, ?, ?, ?)", # (archive_id, file.filename, file.size, file.date, file.crc))
关键说明
pp.Optional(file_metadata)自动跳过开头的元数据块,不影响后续解析pp.OneOrMore(pp.Group(archive))支持多个归档块,Group()将每个归档的结果封装为独立组,方便遍历处理- 优化后的字段匹配规则不再依赖固定后续关键词,提升了解析的健壮性
ZeroOrMore(pp.LineEnd())处理块间的换行和空白,确保解析不受格式换行影响
内容的提问来源于stack exchange,提问作者JK Laiho
相关产品推荐
相关产品推荐

