如何用Python打印指定数量的句子?除NLTK外还有其他方法吗?
替代NLTK的句子提取方法
针对你给出的字符串:
q = "Widows and orphans occur when the first line of a paragraph is the last in a column or page, or when the last line of a paragraph is the first line of a new column or page.The function of a paragraph is to mark a pause, setting the paragraph apart from what precedes it. If a paragraph is preceded by a title or subhead, the indent is superfluous and can therefore be omitted."
除了NLTK,这里提供两种实用的实现方式:
方法一:正则表达式分割
利用示例文本里句子结尾句号后紧跟大写字母的特征,用正则匹配分割边界,代码简单直接:
import re def get_first_n_sentences(text, n): # 匹配句号后紧跟大写字母的位置作为分割点 sentences = re.split(r'(?<=\.)(?=[A-Z])', text) # 取前n个句子并拼接(要单独返回列表就直接 return sentences[:n]) return '. '.join(sentences[:n]) + '.' if n > 0 else '' # 测试取第1个句子 print(get_first_n_sentences(q, 1)) # 测试取前2个句子 print(get_first_n_sentences(q, 2))
注:如果字符串里存在缩写(比如Mr. Smith),这个正则会误分割,但你的示例文本没有这类情况,完全适用。
方法二:手动遍历分割
如果不想依赖正则库,可以手动遍历字符找句子边界,适合需要更精准控制的场景:
def get_first_n_sentences(text, n): sentences = [] current_sentence = [] for idx, char in enumerate(text): current_sentence.append(char) # 遇到句号,检查下一个字符是否是大写字母(或是文本末尾) if char == '.' and (idx == len(text)-1 or text[idx+1].isupper()): sentences.append(''.join(current_sentence)) current_sentence = [] # 取够n个就提前退出 if len(sentences) == n: break # 处理可能没完成的句子(如果n大于实际句子数) if current_sentence and len(sentences) < n: sentences.append(''.join(current_sentence)) return ''.join(sentences[:n]) # 测试示例 print(get_first_n_sentences(q, 1)) print(get_first_n_sentences(q, 2))
这个方法可以根据需求调整判断逻辑,比如处理缩写、特殊标点等复杂情况。
内容的提问来源于stack exchange,提问作者Mohit Narwani
相关产品推荐
相关产品推荐

