You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何提取长字符串中所有关联(L1)的跨行内容并存入列表

实现方案

核心思路

  • 所有独立条目均以x.x.x (Lx)(x为数字)格式作为开头标识,以此为分割规则拆分整个长字符串,即可保留跨多行的同一条目内容不被拆分
  • 拆分后拼接每个条目的前缀和内容,筛选包含(L1)的条目存入结果列表

完整实现代码

import re

teststr = "First Line.............................234" \
          "1.1.0 (L1) TestLine.........................567" \
          "1.1.1 (L1) Second Line.............................587"\
          "Third Line.............................856" \
          "1.1.2 (L2) Fourth Line.............................775"\
          "1.2.7 (L1) Fifth Line.............................262" \
          "1.5.3 (L1) Sixth Line .............................346"\
          "Seventh Line..............................234"

# 匹配条目开头的正则模式:数字.数字.数字 (L数字)
pattern = re.compile(r'(\d+\.\d+\.\d+\s*\(L\d+\))')
# 分割字符串,捕获组会保留匹配到的前缀
split_parts = pattern.split(teststr)

result = []
# 跳过第一个没有前缀的无关内容,从索引1开始每两个元素为一组:[前缀, 内容]
for i in range(1, len(split_parts), 2):
    prefix = split_parts[i]
    content = split_parts[i+1] if i+1 < len(split_parts) else ''
    full_item = prefix + content
    # 筛选L1相关条目
    if '(L1)' in full_item:
        result.append(full_item)

# 输出结果验证
for item in result:
    print(item)
    print('-'*50)

输出结果说明

运行后得到的result列表内容如下:

  • 1.1.0 (L1) TestLine.........................567
  • 1.1.1 (L1) Second Line.............................587Third Line.............................856
  • 1.2.7 (L1) Fifth Line.............................262
  • 1.5.3 (L1) Sixth Line .............................346Seventh Line..............................234
    完全保留了跨多行的同一条目内容,符合需求。

内容的提问来源于stack exchange,提问作者IRezzet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 10:15:00