如何在Python中将三句文档字符串拆分为含标点的句子列表?
拆分文档字符串为完整句子列表的解决方案
修复你的现有代码
你的代码存在两处问题:未初始化存储结果的sentencelist列表,以及找到句子结尾标点后没有重置sentence变量,导致后续句子会拼接前面的内容。修复后的代码如下:
sentencelist = [] sentence = '' docstring = 'The rain in #Spain in 2019, rained "mainly" on the plain.\ There is a nice function to split a string into a list based on a given \ delimiter! Why do I try to do too much?' for character in docstring: sentence += character if character in ('.', '?', '!'): sentencelist.append(sentence.strip()) # 可选:去除句首可能的多余空格 sentence = '' # 重置变量,准备收集下一句 print(sentencelist)
执行后输出:
['The rain in #Spain in 2019, rained "mainly" on the plain.', 'There is a nice function to split a string into a list based on a given delimiter!', 'Why do I try to do too much?']
更高效的正则表达式方案
如果不想用循环遍历字符,Python的re模块可以更简洁地实现需求:通过正向后顾断言拆分字符串,同时保留句子结尾的标点符号。
import re docstring = 'The rain in #Spain in 2019, rained "mainly" on the plain.\ There is a nice function to split a string into a list based on a given \ delimiter! Why do I try to do too much?' # 匹配.?!的位置进行拆分,同时保留标点;过滤空内容并去除多余空格 sentencelist = [s.strip() for s in re.split(r'(?<=[.!?])', docstring) if s.strip()] print(sentencelist)
这个方案更适合处理较长的文本,代码可读性和效率都更高。
内容的提问来源于stack exchange,提问作者Beginner
相关产品推荐
相关产品推荐

