You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从指定文本中提取完整有效句子?

实现思路

步骤拆解

  • 文本预处理:将原始文本按换行符分割为单行,清理每行首尾的空白字符(多余空格、制表符等),同时过滤掉空行。
  • 拆分前缀与内容:对每一行,以第一个冒号为分隔点拆分,分离出前缀标签(如Hobby:、Something (where):)和后续内容,避免误拆分前缀里的特殊字符(比如括号内的内容)。
  • 筛选完整句子:通过规则判断内容是否为完整句子:
    • 基础规则:内容需以大写字母开头、以句号结尾;
    • 补充规则:排除无主谓结构的片段(比如示例中的To Everest mountain,仅为介词短语,无主语和谓语动词),可通过检查内容中是否包含常见谓语动词(如like、want、go等)辅助判断。
  • 收集结果:将所有符合条件的完整句子收集起来,得到最终提取结果。

代码示例(Python)

document = "Hobby: I like going to the mountains.\n Something (where): To Everest mountain.\n\n             The reason: I want to go because I like nature.\n Activities: I'd like to go hiking and admiring the beauty of the nature. "

# 分割并清理行数据
processed_lines = [line.strip() for line in document.split('\n') if line.strip()]
valid_sentences = []

for line in processed_lines:
    if ':' in line:
        # 按第一个冒号拆分,避免前缀含特殊字符时出错
        _, content = line.split(':', 1)
        content = content.strip()
        # 验证完整句子的规则
        if content and content[0].isupper() and content.endswith('.'):
            # 检查是否存在谓语动词,排除非完整片段
            core_verbs = {'like', 'want', 'go', 'admire', 'hike'}
            if any(verb in content.lower() for verb in core_verbs):
                valid_sentences.append(content)

print(valid_sentences)
# 输出结果:
# ['I like going to the mountains.', 'I want to go because I like nature.', "I'd like to go hiking and admiring the beauty of the nature."]

内容的提问来源于stack exchange,提问作者stephsmith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 18:05:48