如何用Python正则表达式拆分文本并保留句末句号?
问题描述
我需要将文本:
"Hi stackoverflow.I need to split these sentences.By full stop."
拆分为如下格式的列表:
["Hi stackoverflow.", "I need to split these sentences.", "By full stop."]
我尝试了以下Python代码:
import regex as re text = "Hi stackoverflow.I need to split these sentences.By full stop." re.findall(".*?\.[a-zA-Z0-9].*?", text)
但执行后返回结果为['Hi stackoverflow.I', ' need to split these sentences.B'],不符合预期,请问如何正确实现需求?
解决方案
方法一:正则直接匹配完整句子
你原正则的问题在于,会把句号和后续的首字母连在一起匹配,导致截取的内容不符合预期。针对你的文本结构(句子间以「句号+大写字母」衔接,首句以小写开头),可以用正则匹配两种句式:
import re text = "Hi stackoverflow.I need to split these sentences.By full stop." sentences = re.findall(r'^[^.]+\.|[A-Z][^.]+\.', text) print(sentences)
执行后输出:
["Hi stackoverflow.", "I need to split these sentences.", "By full stop."]
方法二:预处理文本后拆分
如果文本规则固定(句号后紧跟下一句的首字母),也可以先给句号后补空格,再拆分句子,逻辑更直观:
import re text = "Hi stackoverflow.I need to split these sentences.By full stop." # 在句号与后续字母之间插入空格 fixed_text = re.sub(r'\.([A-Za-z])', r'. \1', text) # 按". "拆分后,给每个片段补回句号 sentences = [f"{s}." for s in fixed_text.split('. ') if s] print(sentences)
内容的提问来源于stack exchange,提问作者ranni rabadi
相关产品推荐
相关产品推荐

