You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分多行文本字段并在指定分段保留换行符后解析为键值结构

实现思路

核心逻辑

  • 提前维护需要保留换行的字段集合,后续新增同类字段仅需更新该集合,无需修改整体解析逻辑
  • 逐行扫描文本,通过正则识别键值对的起始边界,普通字段直接取单行值,指定字段持续拼接后续行直到下一个新键出现
  • 用有序字典存储解析结果,天然支持按插入顺序/自定义顺序输出

具体解析步骤

  1. 预定义保留换行的字段集合,示例:keep_newline_fields = {"Comment"}
  2. 初始化三个变量:
    • current_key:记录当前正在处理的键名,初始为空
    • current_value:记录当前键对应的值,初始为空
    • 有序字典result存储最终结果
  3. 逐行遍历输入文本:
    • 用正则^([a-zA-Z0-9_]+):\s(.*)匹配当前行是否为新键的起始行
    • 若匹配到新键:
      • 若current_key非空,先将current_key和处理后的current_value存入result
      • 更新current_key为匹配到的键名,current_value为匹配到的该行值内容
    • 若未匹配到新键:
      • 若current_key属于keep_newline_fields,直接将当前行拼接到current_value末尾,保留换行符
      • 否则直接跳过当前行(普通字段仅取冒号所在行的内容)
  4. 遍历结束后,将最后一组current_key和current_value存入result

示例代码(Python)

import re
from collections import OrderedDict

def parse_multi_line_text(input_text, keep_newline_fields):
    # 3.7+ 版本Python可直接用dict代替OrderedDict,默认保留插入顺序
    result = OrderedDict()
    current_key = None
    current_value = []
    
    key_pattern = re.compile(r'^([a-zA-Z0-9_]+):\s(.*)')
    
    for line in input_text.splitlines(keepends=True):
        match = key_pattern.match(line)
        if match:
            # 遇到新键,先保存上一个键值对
            if current_key is not None:
                if current_key in keep_newline_fields:
                    result[current_key] = ''.join(current_value).rstrip('\n')
                else:
                    result[current_key] = ''.join(current_value).strip()
            # 初始化新键的内容
            current_key = match.group(1)
            current_value = [match.group(2)]
        else:
            # 非新键行,仅当当前键需要保留换行时拼接
            if current_key in keep_newline_fields:
                current_value.append(line)
    
    # 保存最后一个键值对
    if current_key is not None:
        if current_key in keep_newline_fields:
            result[current_key] = ''.join(current_value).rstrip('\n')
        else:
            result[current_key] = ''.join(current_value).strip()
    
    return result

# 测试用例
test_input = """title1: 11.11.2021
title2: documentation
title3: right
Comment: Here 
I still have no 
idea how to keep the 
line breaks

title4: something
"""

keep_fields = {"Comment"}
parsed_data = parse_multi_line_text(test_input, keep_fields)

# 输出测试
for k, v in parsed_data.items():
    print(f"{k}: {repr(v)}")

运行后输出的Comment值为'Here \nI still have no \nidea how to keep the \nline breaks',完整保留了所有换行符,其余字段为普通单行值。

扩展方案

如果需要支持更复杂的结构化存储,可直接将解析后的有序字典序列化为JSON格式存储,多行列字段的换行符会自动转义为\n,读取时还原即可,无需额外处理。


内容的提问来源于stack exchange,提问作者Tim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 16:15:04