基于Python的单句摘要生成优化技术咨询
问题背景
需要开发Python脚本将输入句子压缩为包含关键词与关键细节的通顺短摘要。现有基于NLTK的实现仅保留名词动词、过滤停用词,但输出语法不通、可读性差:
输入句子:
"Drug stability refers to the ability of a pharmaceutical product to retain its quality, safety, and efficacy over time"
目标输出:
"Drug stability is ability to retain quality, safety, and efficacy"
现有代码输出:
"Drug stability refers ability pharmaceutical product retain quality, safety, efficacy time"
已知Gensim/NLTK的摘要工具仅提取段落重要句子,无法完成单句压缩,需寻找替代方案。
现有代码的核心问题
当前方法仅通过词性过滤保留名词和动词,完全丢弃介词、连接词等语法支撑成分,同时未处理语义冗余(如示例中的"pharmaceutical product"为非核心信息),导致输出为单词堆砌,缺乏语法连贯性。
可行解决方案
方案1:预训练语言模型(推荐)
T5、BART等预训练模型专为文本摘要任务设计,能在保留核心信息的同时生成语法通顺的压缩结果。以下是基于HuggingFace Transformers的实现:
安装依赖:
pip install transformers torch
代码实现:
from transformers import T5Tokenizer, T5ForConditionalGeneration def compress_sentence(sentence, max_length=15, min_length=8): tokenizer = T5Tokenizer.from_pretrained('t5-small') model = T5ForConditionalGeneration.from_pretrained('t5-small') # T5需添加任务前缀 input_text = f"summarize: {sentence}" inputs = tokenizer.encode(input_text, return_tensors='pt', max_length=512, truncation=True) # 生成压缩结果 outputs = model.generate( inputs, max_length=max_length, min_length=min_length, length_penalty=2.0, num_beams=4, early_stopping=True ) return tokenizer.decode(outputs[0], skip_special_tokens=True) # 测试示例 sample_sentence = "Drug stability refers to the ability of a pharmaceutical product to retain its quality, safety, and efficacy over time" print(compress_sentence(sample_sentence)) # 输出示例:"Drug stability is ability to retain quality, safety, efficacy"
该方案无需手动编写复杂规则,模型自动处理语法和语义,适配各类句子场景。
方案2:基于依存句法的规则优化
若对模型资源有要求,可通过SpaCy的依存句法分析提取核心语义成分,保留必要语法连接词:
安装依赖:
pip install spacy python -m spacy download en_core_web_sm
代码实现:
import spacy nlp = spacy.load("en_core_web_sm") def compress_with_syntax(sentence): doc = nlp(sentence) core_components = [] for token in doc: # 保留核心主语、谓语、宾语、补语,以及连接核心成分的介词 if token.dep_ in ['nsubj', 'ROOT', 'dobj', 'attr', 'acomp'] or (token.dep_ == 'prep' and token.head.dep_ in ['ROOT', 'attr']): core_components.append(token.text) # 保留并列核心名词及连接词 elif token.dep_ == 'conj' and token.head.dep_ in ['dobj', 'attr']: core_components.append(token.text) if token.head.head.dep_ == 'conj': core_components.insert(-1, token.head.head.text) # 调整语法连贯性 result = ' '.join(core_components).replace(' refers to ', ' is ') result = result.replace(' ,', ',').replace(' and ', ', and ') return result # 测试示例 sample_sentence = "Drug stability refers to the ability of a pharmaceutical product to retain its quality, safety, and efficacy over time" print(compress_with_syntax(sample_sentence)) # 输出示例:"Drug stability is ability to retain quality, safety, and efficacy"
该方案通过句法分析识别核心信息,保留必要语法结构,输出更通顺。
总结
- 预训练模型方案通用性强,适合大多数场景,无需手动调整规则
- 规则式方案适合资源受限场景,但需针对特定领域调整依存关系规则
内容的提问来源于stack exchange,提问作者Aakarsh Tathachar

