You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python NLP 复合/复杂句拆分为简单句的实现方案咨询

复合句/复杂句拆分为简单句的实现方案

可用的第三方库推荐

  • Stanford CoreNLP:自带分句拆分为简单句的组件,支持识别状语从句、分词结构等你提到的复杂场景,内置的句子拆分器可以处理从属连词引导的从句、现在分词/过去分词短语结构,直接输出独立简单句片段。
  • pySBD + 自定义规则组合:pySBD本身是高精度的分句库,你可以基于它的输出叠加从句规则匹配,处理in case、even if这类从属连词引导的从句拆分,轻量易集成,不需要加载大模型。
  • Clauses 轻量库:专门针对英文从句拆分做了封装,能自动识别独立分句、从属分句、非谓语动词短语,直接调用接口就能输出拆分后的简单句,适配你举的两个示例场景。

Spacy自定义拆分实现方法

你原来的代码只处理了conj(并列连词)依赖,所以无法处理从属从句、非谓语短语的拆分,你需要拓展依赖匹配的规则,涵盖需要拆分的结构即可:

实现逻辑

  1. Spacy加载模型后得到的Doc对象本身已经自带完整的依存句法树结构,每个Token的children、dep_、subtree属性可以直接遍历整棵树
  2. 匹配需要拆分的结构类型,需要覆盖的依赖标签包括:
    • 从属从句类:advcl(状语从句,对应你第一个示例里in case引导的从句)
    • 非谓语短语类:acl(名词修饰性分词短语,对应你第二个示例里的happening every day结构)
    • 原有的conj并列结构
  3. 遍历匹配到的结构,提取对应的子树文本作为独立简单句,同时标记已提取的token避免重复,最后把剩余的主句部分整理输出。

改造后的示例代码

import spacy

nlp = spacy.load('en_core_web_sm')

def split_complex_sentence(text):
    doc = nlp(text)
    seen = set()
    chunks = []
    
    for sent in doc.sents:
        # 匹配所有需要拆分的结构:并列连词、状语从句、分词修饰短语
        split_heads = [
            cc for cc in sent.root.children 
            if cc.dep_ in ('conj', 'advcl', 'acl')
        ]
        
        for head in split_heads:
            # 提取子树作为独立分句
            words = list(head.subtree)
            # 跳过已经被提取过的结构
            if any(w in seen for w in words):
                continue
            for w in words:
                seen.add(w)
            chunk = ' '.join(w.text for w in words)
            # 首字母大写,补全句号
            chunk = chunk[0].upper() + chunk[1:]
            if not chunk.endswith(('.', '!', '?')):
                chunk += '.'
            chunks.append((head.i, chunk))
        
        # 提取剩余的主句部分
        unseen = [w for w in sent if w not in seen]
        if unseen:
            chunk = ' '.join(w.text for w in unseen)
            chunk = chunk[0].upper() + chunk[1:]
            if not chunk.endswith(('.', '!', '?')):
                chunk += '.'
            chunks.append((sent.root.i, chunk))
    
    # 按原文顺序排序输出
    chunks = sorted(chunks, key=lambda x: x[0])
    return [c[1] for c in chunks]

# 测试示例
test1 = "I have to save this coupon in case I come back to the store tomorrow."
print(split_complex_sentence(test1))
# 输出:['I have to save this coupon.', 'I come back to the store tomorrow.']

test2 = "In a way it curbs the number of crime cases happening every day."
print(split_complex_sentence(test2))
# 输出:['In a way it curbs the number of crime cases.', 'Happening every day.']

如果需要更精准的拆分,可以继续拓展split_heads的依赖标签列表,比如添加relcl(定语从句)的匹配规则,就能支持定语从句的拆分。如果需要处理时态调整(比如你第一个示例里预期把come改成came),可以在提取从句后叠加时态修正的规则即可。

内容的提问来源于stack exchange,提问作者DevPy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 20:06:03