如何拆分多行文本字段并在指定分段保留换行符后解析为键值结构
实现思路
核心逻辑
- 提前维护需要保留换行的字段集合,后续新增同类字段仅需更新该集合,无需修改整体解析逻辑
- 逐行扫描文本,通过正则识别键值对的起始边界,普通字段直接取单行值,指定字段持续拼接后续行直到下一个新键出现
- 用有序字典存储解析结果,天然支持按插入顺序/自定义顺序输出
具体解析步骤
- 预定义保留换行的字段集合,示例:
keep_newline_fields = {"Comment"} - 初始化三个变量:
current_key:记录当前正在处理的键名,初始为空current_value:记录当前键对应的值,初始为空- 有序字典
result存储最终结果
- 逐行遍历输入文本:
- 用正则
^([a-zA-Z0-9_]+):\s(.*)匹配当前行是否为新键的起始行 - 若匹配到新键:
- 若
current_key非空,先将current_key和处理后的current_value存入result - 更新
current_key为匹配到的键名,current_value为匹配到的该行值内容
- 若
- 若未匹配到新键:
- 若
current_key属于keep_newline_fields,直接将当前行拼接到current_value末尾,保留换行符 - 否则直接跳过当前行(普通字段仅取冒号所在行的内容)
- 若
- 用正则
- 遍历结束后,将最后一组
current_key和current_value存入result
示例代码(Python)
import re from collections import OrderedDict def parse_multi_line_text(input_text, keep_newline_fields): # 3.7+ 版本Python可直接用dict代替OrderedDict,默认保留插入顺序 result = OrderedDict() current_key = None current_value = [] key_pattern = re.compile(r'^([a-zA-Z0-9_]+):\s(.*)') for line in input_text.splitlines(keepends=True): match = key_pattern.match(line) if match: # 遇到新键,先保存上一个键值对 if current_key is not None: if current_key in keep_newline_fields: result[current_key] = ''.join(current_value).rstrip('\n') else: result[current_key] = ''.join(current_value).strip() # 初始化新键的内容 current_key = match.group(1) current_value = [match.group(2)] else: # 非新键行,仅当当前键需要保留换行时拼接 if current_key in keep_newline_fields: current_value.append(line) # 保存最后一个键值对 if current_key is not None: if current_key in keep_newline_fields: result[current_key] = ''.join(current_value).rstrip('\n') else: result[current_key] = ''.join(current_value).strip() return result # 测试用例 test_input = """title1: 11.11.2021 title2: documentation title3: right Comment: Here I still have no idea how to keep the line breaks title4: something """ keep_fields = {"Comment"} parsed_data = parse_multi_line_text(test_input, keep_fields) # 输出测试 for k, v in parsed_data.items(): print(f"{k}: {repr(v)}")
运行后输出的Comment值为'Here \nI still have no \nidea how to keep the \nline breaks',完整保留了所有换行符,其余字段为普通单行值。
扩展方案
如果需要支持更复杂的结构化存储,可直接将解析后的有序字典序列化为JSON格式存储,多行列字段的换行符会自动转义为\n,读取时还原即可,无需额外处理。
内容的提问来源于stack exchange,提问作者Tim
相关产品推荐
相关产品推荐

