You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python循环构建正则处理多页文本:首尾截断失效问题求助

处理文本列表并合并的解决方案

核心思路

  • 单独处理首段:只保留marcacao_comeco之后的内容
  • 单独处理末段:只保留marcacao_fim之前的内容
  • 中间段落直接原样保留
  • 最后把所有处理后的段落拼接成完整字符串

示例代码

假设你的文本列表是这样的:

text_list = [
    "Some useless text before marcacao_comeco girl named Lily who lived in a small village at...",
    "She loved walking in the forest every morning, picking berries and listening to birds sing.",
    "One day she encountered a big brown bear, but marcacao_fim some irrelevant text after this"
]

对应的处理代码:

import re

text_list = [
    "Some useless text before marcacao_comeco girl named Lily who lived in a small village at...",
    "She loved walking in the forest every morning, picking berries and listening to birds sing.",
    "One day she encountered a big brown bear, but marcacao_fim some irrelevant text after this"
]

# 处理首段:剔除marcacao_comeco之前的所有内容
first_part = re.sub(r'^.*marcacao_comeco', '', text_list[0]).strip() if text_list else ""

# 处理末段和中间段落
if len(text_list) > 1:
    last_part = re.sub(r'marcacao_fim.*$', '', text_list[-1]).strip()
    middle_parts = text_list[1:-1]
else:
    # 列表只有一段的情况,同时处理首尾标记
    last_part = re.sub(r'marcacao_fim.*$', '', text_list[0]).strip() if text_list else ""
    middle_parts = []

# 拼接所有部分得到结果
final_text = ' '.join([first_part] + middle_parts + [last_part])
print(final_text)

为什么之前的re.sub没生效?

大概率是没给正则加上定位符:

  • 处理首段时要加^匹配字符串开头,确保只删掉开头到marcacao_comeco的内容,而不是所有出现marcacao_comeco的地方
  • 处理末段时要加$匹配字符串结尾,确保只删掉marcacao_fim到结尾的内容

运行上面的代码后,输出就会符合你的预期:以girl named Lily who lived in a small village at ...开头,以... the bear did not attack her.结尾。

内容的提问来源于stack exchange,提问作者jvqp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 01:55:06