You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jupyter Notebook中使用sumy文本摘要遇UnpicklingError报错求助

解决Sumy文本摘要中的UnpicklingError问题

问题原因

这个UnpicklingError: global 'copy_reg._reconstructor' is forbidden错误,是因为新版NLTK(3.9+)引入了严格的RestrictedUnpickler机制,禁止加载pickle文件时调用copy_reg._reconstructor,而Sumy依赖的旧版NLTK punkt分词模型pickle文件正好需要这个函数。

解决方案

方案1:降级NLTK到兼容版本

这是最直接的解决方法:

  • 卸载当前NLTK版本,安装兼容的3.8.1版本:
    pip uninstall -y nltk
    pip install nltk==3.8.1
    
  • 在Jupyter Notebook中重新下载punkt资源:
    import nltk
    nltk.download('punkt')
    
  • 之后运行你原来的Sumy代码即可正常工作。

方案2:自定义Tokenizer绕过Sumy的默认实现(无需降级NLTK)

如果不想降级NLTK,可以自定义一个Tokenizer,直接调用NLTK的原生分词函数,绕开Sumy加载pickle文件的逻辑:

import nltk
nltk.download('punkt')
nltk.download('punkt_tab')

from sumy.parsers.plaintext import PlaintextParser
from sumy.summarizers.lsa import LsaSummarizer
from sumy.nlp.stemmers import Stemmer
from sumy.utils import get_stop_words

# 自定义Tokenizer,直接调用NLTK的分词接口
class CustomTokenizer:
    def __init__(self, language):
        self.language = language

    def tokenize_sentences(self, text):
        return nltk.sent_tokenize(text, language=self.language)
    
    def tokenize_words(self, sentence):
        return nltk.word_tokenize(sentence, language=self.language)

# 待摘要的文本
text = """
Your long text here...
"""

# 使用自定义Tokenizer创建解析器
parser = PlaintextParser.from_string(text, CustomTokenizer("english"))

# 初始化摘要器
stemmer = Stemmer("english")
summarizer = LsaSummarizer(stemmer)
summarizer.stop_words = get_stop_words("english")

# 生成3句摘要
summary = summarizer(parser.document, 3)

# 打印结果
for sentence in summary:
    print(sentence)

这个自定义Tokenizer完全避开了Sumy默认加载pickle的流程,直接用NLTK的分词函数,适配新版NLTK的限制。

内容的提问来源于stack exchange,提问作者john doe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 11:27:43