You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式拆分文本并保留句末句号?

问题描述

我需要将文本:

"Hi stackoverflow.I need to split these sentences.By full stop."
拆分为如下格式的列表:
["Hi stackoverflow.", "I need to split these sentences.", "By full stop."]
我尝试了以下Python代码:

import regex as re
text = "Hi stackoverflow.I need to split these sentences.By full stop."
re.findall(".*?\.[a-zA-Z0-9].*?", text)

但执行后返回结果为['Hi stackoverflow.I', ' need to split these sentences.B'],不符合预期,请问如何正确实现需求?


解决方案

方法一:正则直接匹配完整句子

你原正则的问题在于,会把句号和后续的首字母连在一起匹配,导致截取的内容不符合预期。针对你的文本结构(句子间以「句号+大写字母」衔接,首句以小写开头),可以用正则匹配两种句式:

import re
text = "Hi stackoverflow.I need to split these sentences.By full stop."
sentences = re.findall(r'^[^.]+\.|[A-Z][^.]+\.', text)
print(sentences)

执行后输出:

["Hi stackoverflow.", "I need to split these sentences.", "By full stop."]

方法二:预处理文本后拆分

如果文本规则固定(句号后紧跟下一句的首字母),也可以先给句号后补空格,再拆分句子,逻辑更直观:

import re
text = "Hi stackoverflow.I need to split these sentences.By full stop."
# 在句号与后续字母之间插入空格
fixed_text = re.sub(r'\.([A-Za-z])', r'. \1', text)
# 按". "拆分后,给每个片段补回句号
sentences = [f"{s}." for s in fixed_text.split('. ') if s]
print(sentences)

内容的提问来源于stack exchange,提问作者ranni rabadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 06:50:29