You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SUMY文本摘要工具LuhnSummarizer失效:返回原文而非摘要

Troubleshooting SUMY's LuhnSummarizer Returning Full Text Instead of Summary

Got it, let’s figure out why your LuhnSummarizer is just spitting back the original text instead of a condensed summary. I’ve hit similar snags with SUMY before, so here are the most likely fixes:

Check Sentence Structure & Count

First up, LuhnSummarizer relies on the parser correctly splitting your text into individual sentences. If your textA lacks proper sentence-ending punctuation (like periods, question marks), the PlaintextParser might treat the entire block as a single sentence. When you ask for 10 sentences, it just returns that one big "sentence" (your whole text).

Test this quickly by adding a print statement in your function:

parser = PlaintextParser.from_string(text, Tokenizer(LANGUAGE))
print(f"Total sentences detected: {len(parser.document.sentences)}")

If the output is 1, you’ll need to fix your text’s sentence separators. Alternatively, pre-process the text with a more robust sentence tokenizer (like NLTK’s sent_tokenize) before passing it to SUMY.

Bind the Stemmer to the Summarizer

Your code initializes a Stemmer but never attaches it to the LuhnSummarizer. Luhn’s algorithm uses stemming to identify important, recurring keywords—without it, the summarizer can’t tell which sentences are more significant, so it defaults to returning everything.

Fix this by passing the stemmer when creating the summarizer:

summarizer = LuhnSummarizer(stemmer)

Adjust the Summary Length Parameter

You’re asking for 10 sentences with summarizer(parser.document, 10). If your original text has 10 or fewer sentences total, the summarizer will just return all of them. Instead of a fixed number, try using a ratio to get a percentage of the original text:

# Returns 20% of the original sentences
summary_sentences = summarizer(parser.document, ratio=0.2)

Or add a check in your function to cap the requested sentence count at the total available.

Verify Your SUMY Version

Older versions of SUMY sometimes had API quirks that could cause this behavior. Run print(sumy.__version__) to check—if it’s not the latest, upgrade with pip install --upgrade sumy.

Modified Working Example

Here’s your code with these fixes applied:

from sumy.parsers.plaintext import PlaintextParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.luhn import LuhnSummarizer
from sumy.nlp.stemmers import Stemmer
from sumy.utils import get_stop_words

LANGUAGE = "english"
stemmer = Stemmer(LANGUAGE)

def get_luhn_summary(text, target_sentences=10):
    parser = PlaintextParser.from_string(text, Tokenizer(LANGUAGE))
    total_sentences = len(parser.document.sentences)
    
    # Don't ask for more sentences than exist
    target = min(target_sentences, total_sentences)
    if target == total_sentences:
        return [str(sentence) for sentence in parser.document.sentences]
    
    summarizer = LuhnSummarizer(stemmer)
    summarizer.stop_words = get_stop_words(LANGUAGE)
    
    summary = summarizer(parser.document, target)
    return [str(sentence) for sentence in summary]

# Usage
summaryA_luhn = get_luhn_summary(textA)

内容的提问来源于stack exchange,提问作者jokol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 10:48:17