如何用Pyparsing处理带引号换行拼接的键值对解析?
用pyparsing优雅处理引号字符串拼接的键值对
好问题!这种相邻引号字符串直接拼接的场景完全不用提前做文本替换,pyparsing可以直接在解析阶段优雅处理。咱们只需要调整一下现有的解析规则,让value能够匹配一个或多个连续的带引号字符串,再把它们的内容合并起来就行。
调整后的解析规则
先看修改后的完整代码(我假设你的key是字母数字+下划线组成的,如果你的key规则有特殊要求,直接替换key的定义就行):
from pyparsing import Literal, QuotedString, Group, OneOrMore, Word, alphanums # 定义key(可根据实际需求扩展,比如支持连字符等) key = Word(alphanums + "_") eq = Literal('=').suppress() # 统一用支持多行和转义双引号的QuotedString(原来的v1和v2可以合并,multiline=True也兼容单行场景) quoted_str = QuotedString('"', multiline=True, escQuote='""') # 让value匹配一个或多个连续的带引号字符串,然后合并内容 # 如果拼接字符串之间需要保留空格就用' '.join,不需要则用''.join value = Group(OneOrMore(quoted_str)).setParseAction( lambda tokens: ' '.join(tokens[0]) )("value") kv = Group(key + eq + value)("key_value")
测试你的示例场景
用你给出的key5示例测试一下:
test_input = 'key5="some more text that is so long that the authors who serialized it to a file thought it" "would be a good idea to to concatenate strings this way"' result = kv.parseString(test_input) print(result.dump())
输出结果会把两个引号里的内容合并成一个完整的value:
[['key5', 'some more text that is so long that the authors who serialized it to a file thought it would be a good idea to to concatenate strings this way']] - key_value: ['key5', 'some more text that is so long that the authors who serialized it to a file thought it would be a good idea to to concatenate strings this way'] - value: 'some more text that is so long that the authors who serialized it to a file thought it would be a good idea to to concatenate strings this way'
方案优势
- 无需预处理:直接在解析阶段完成拼接,不用提前修改原始文本,保持代码简洁性
- 兼容原有功能:完美保留了你原来支持的多行文本和
""转义双引号的逻辑 - 灵活可调:如果你的场景中拼接字符串不需要保留中间空格,只需要把
' '.join改成''.join即可
内容的提问来源于stack exchange,提问作者Hank
相关产品推荐
相关产品推荐

