You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python打印指定数量的句子?除NLTK外还有其他方法吗?

替代NLTK的句子提取方法

针对你给出的字符串:

q = "Widows and orphans occur when the first line of a paragraph is the last in a column or page, or when the last line of a paragraph is the first line of a new column or page.The function of a paragraph is to mark a pause, setting the paragraph apart from what precedes it. If a paragraph is preceded by a title or subhead, the indent is superfluous and can therefore be omitted."

除了NLTK,这里提供两种实用的实现方式:

方法一:正则表达式分割

利用示例文本里句子结尾句号后紧跟大写字母的特征,用正则匹配分割边界,代码简单直接:

import re

def get_first_n_sentences(text, n):
    # 匹配句号后紧跟大写字母的位置作为分割点
    sentences = re.split(r'(?<=\.)(?=[A-Z])', text)
    # 取前n个句子并拼接(要单独返回列表就直接 return sentences[:n])
    return '. '.join(sentences[:n]) + '.' if n > 0 else ''

# 测试取第1个句子
print(get_first_n_sentences(q, 1))
# 测试取前2个句子
print(get_first_n_sentences(q, 2))

注:如果字符串里存在缩写(比如Mr. Smith),这个正则会误分割,但你的示例文本没有这类情况,完全适用。

方法二:手动遍历分割

如果不想依赖正则库,可以手动遍历字符找句子边界,适合需要更精准控制的场景:

def get_first_n_sentences(text, n):
    sentences = []
    current_sentence = []
    for idx, char in enumerate(text):
        current_sentence.append(char)
        # 遇到句号,检查下一个字符是否是大写字母(或是文本末尾)
        if char == '.' and (idx == len(text)-1 or text[idx+1].isupper()):
            sentences.append(''.join(current_sentence))
            current_sentence = []
            # 取够n个就提前退出
            if len(sentences) == n:
                break
    # 处理可能没完成的句子(如果n大于实际句子数)
    if current_sentence and len(sentences) < n:
        sentences.append(''.join(current_sentence))
    return ''.join(sentences[:n])

# 测试示例
print(get_first_n_sentences(q, 1))
print(get_first_n_sentences(q, 2))

这个方法可以根据需求调整判断逻辑,比如处理缩写、特殊标点等复杂情况。

内容的提问来源于stack exchange,提问作者Mohit Narwani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 04:45:39