Python循环构建正则处理多页文本:首尾截断失效问题求助
处理文本列表并合并的解决方案
核心思路
- 单独处理首段:只保留
marcacao_comeco之后的内容 - 单独处理末段:只保留
marcacao_fim之前的内容 - 中间段落直接原样保留
- 最后把所有处理后的段落拼接成完整字符串
示例代码
假设你的文本列表是这样的:
text_list = [ "Some useless text before marcacao_comeco girl named Lily who lived in a small village at...", "She loved walking in the forest every morning, picking berries and listening to birds sing.", "One day she encountered a big brown bear, but marcacao_fim some irrelevant text after this" ]
对应的处理代码:
import re text_list = [ "Some useless text before marcacao_comeco girl named Lily who lived in a small village at...", "She loved walking in the forest every morning, picking berries and listening to birds sing.", "One day she encountered a big brown bear, but marcacao_fim some irrelevant text after this" ] # 处理首段:剔除marcacao_comeco之前的所有内容 first_part = re.sub(r'^.*marcacao_comeco', '', text_list[0]).strip() if text_list else "" # 处理末段和中间段落 if len(text_list) > 1: last_part = re.sub(r'marcacao_fim.*$', '', text_list[-1]).strip() middle_parts = text_list[1:-1] else: # 列表只有一段的情况,同时处理首尾标记 last_part = re.sub(r'marcacao_fim.*$', '', text_list[0]).strip() if text_list else "" middle_parts = [] # 拼接所有部分得到结果 final_text = ' '.join([first_part] + middle_parts + [last_part]) print(final_text)
为什么之前的re.sub没生效?
大概率是没给正则加上定位符:
- 处理首段时要加
^匹配字符串开头,确保只删掉开头到marcacao_comeco的内容,而不是所有出现marcacao_comeco的地方 - 处理末段时要加
$匹配字符串结尾,确保只删掉marcacao_fim到结尾的内容
运行上面的代码后,输出就会符合你的预期:以girl named Lily who lived in a small village at ...开头,以... the bear did not attack her.结尾。
内容的提问来源于stack exchange,提问作者jvqp
相关产品推荐
相关产品推荐

