You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python的单句摘要生成优化技术咨询

单句压缩摘要解决方案

问题背景

需要开发Python脚本将输入句子压缩为包含关键词与关键细节的通顺短摘要。现有基于NLTK的实现仅保留名词动词、过滤停用词,但输出语法不通、可读性差:

输入句子:

"Drug stability refers to the ability of a pharmaceutical product to retain its quality, safety, and efficacy over time"

目标输出:

"Drug stability is ability to retain quality, safety, and efficacy"

现有代码输出:

"Drug stability refers ability pharmaceutical product retain quality, safety, efficacy time"

已知Gensim/NLTK的摘要工具仅提取段落重要句子,无法完成单句压缩,需寻找替代方案。

现有代码的核心问题

当前方法仅通过词性过滤保留名词和动词,完全丢弃介词、连接词等语法支撑成分,同时未处理语义冗余(如示例中的"pharmaceutical product"为非核心信息),导致输出为单词堆砌,缺乏语法连贯性。

可行解决方案

方案1:预训练语言模型(推荐)

T5、BART等预训练模型专为文本摘要任务设计,能在保留核心信息的同时生成语法通顺的压缩结果。以下是基于HuggingFace Transformers的实现:

安装依赖:

pip install transformers torch

代码实现:

from transformers import T5Tokenizer, T5ForConditionalGeneration

def compress_sentence(sentence, max_length=15, min_length=8):
    tokenizer = T5Tokenizer.from_pretrained('t5-small')
    model = T5ForConditionalGeneration.from_pretrained('t5-small')
    
    # T5需添加任务前缀
    input_text = f"summarize: {sentence}"
    inputs = tokenizer.encode(input_text, return_tensors='pt', max_length=512, truncation=True)
    
    # 生成压缩结果
    outputs = model.generate(
        inputs,
        max_length=max_length,
        min_length=min_length,
        length_penalty=2.0,
        num_beams=4,
        early_stopping=True
    )
    
    return tokenizer.decode(outputs[0], skip_special_tokens=True)

# 测试示例
sample_sentence = "Drug stability refers to the ability of a pharmaceutical product to retain its quality, safety, and efficacy over time"
print(compress_sentence(sample_sentence))
# 输出示例:"Drug stability is ability to retain quality, safety, efficacy"

该方案无需手动编写复杂规则,模型自动处理语法和语义,适配各类句子场景。

方案2:基于依存句法的规则优化

若对模型资源有要求,可通过SpaCy的依存句法分析提取核心语义成分,保留必要语法连接词:

安装依赖:

pip install spacy
python -m spacy download en_core_web_sm

代码实现:

import spacy

nlp = spacy.load("en_core_web_sm")

def compress_with_syntax(sentence):
    doc = nlp(sentence)
    core_components = []
    
    for token in doc:
        # 保留核心主语、谓语、宾语、补语,以及连接核心成分的介词
        if token.dep_ in ['nsubj', 'ROOT', 'dobj', 'attr', 'acomp'] or (token.dep_ == 'prep' and token.head.dep_ in ['ROOT', 'attr']):
            core_components.append(token.text)
        # 保留并列核心名词及连接词
        elif token.dep_ == 'conj' and token.head.dep_ in ['dobj', 'attr']:
            core_components.append(token.text)
            if token.head.head.dep_ == 'conj':
                core_components.insert(-1, token.head.head.text)
    
    # 调整语法连贯性
    result = ' '.join(core_components).replace(' refers to ', ' is ')
    result = result.replace(' ,', ',').replace(' and ', ', and ')
    return result

# 测试示例
sample_sentence = "Drug stability refers to the ability of a pharmaceutical product to retain its quality, safety, and efficacy over time"
print(compress_with_syntax(sample_sentence))
# 输出示例:"Drug stability is ability to retain quality, safety, and efficacy"

该方案通过句法分析识别核心信息,保留必要语法结构,输出更通顺。

总结

  • 预训练模型方案通用性强,适合大多数场景,无需手动调整规则
  • 规则式方案适合资源受限场景,但需针对特定领域调整依存关系规则

内容的提问来源于stack exchange,提问作者Aakarsh Tathachar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 10:20:06