You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决pke调用Spacy时文本超限及分句相关报错问题

解决pke处理长文本时的Spacy相关报错问题

问题描述

使用pke的extractor.load_document()处理长文本时触发报错:

ValueError: [E088] Text of length 1717453 exceeds maximum of 1000000. The parser and NER models require roughly 1GB of temporary memory per 100,000 characters in the input. This means long texts may cause memory allocation errors. If you're not using the parser or NER, it's probably safe to increase the nlp.max_length limit. The limit is in number of characters, so you can check whether your inputs are too long by checking len(text).

尝试两种方案后遇到新问题:

  • 手动提高nlp.max_length未解决根本问题
  • 传入预处理的Spacy Doc对象时,触发句子边界未设置的报错;添加sentencizer组件后又无法提取关键词

原核心代码:

def pke_topicrank(text):
    # initialize keyphrase extraction model, here TopicRank
    extractor = pke.unsupervised.TopicRank()

    # load the content of the document, here document is expected to be a simple 
    # test string and preprocessing is carried out using spacy
    
    #docs = list(nlp.pipe(text, batch_size=1000))
    extractor.load_document(input=text, language="en", \
                            normalization=None)

    # keyphrase candidate selection, in the case of TopicRank: sequences of nouns
    # and adjectives (i.e. `(Noun|Adj)*`)
    pos = {'NOUN', 'PROPN', 'ADJ'}
    extractor.candidate_selection(pos=pos)
    #extractor.candidate_selection()
    
    #grammar selection
    extractor.grammar_selection(grammar="NP: {<ADJ>*<NOUN|PROPN>+}")

    # candidate weighting, in the case of TopicRank: using a random walk algorithm
    extractor.candidate_weighting(threshold=0.74, method='average')

    # N-best selection, keyphrases contains the 10 highest scored candidates as
    # (keyphrase, score) tuples
    keyphrases = extractor.get_n_best(n=10, redundancy_removal=True, stemming=True)
    keyphrases = ', '.join(set([candidate for candidate, weight in keyphrases]))
    return keyphrases

解决方案

步骤1:正确配置Spacy Pipeline

必须保留词性标注(tagger)组件(TopicRank依赖词性筛选候选词),添加sentencizer处理句子边界,同时调整max_length适配长文本:

import spacy
# 优先启用GPU加速
activated = spacy.prefer_gpu()
# 加载模型时排除不需要的parser和ner,保留tagger
nlp = spacy.load('en_core_web_sm', exclude=['parser', 'ner'])
# 添加sentencizer组件用于拆分句子
nlp.add_pipe('sentencizer')
# 设置足够大的max_length,确保覆盖输入文本长度
nlp.max_length = 2000000  # 可根据实际文本长度调整

步骤2:修改pke加载文档逻辑

先通过配置好的Spacy处理文本得到Doc对象,再传入pke,同时确保参数匹配:

def pke_topicrank(text):
    extractor = pke.unsupervised.TopicRank()
    
    # 用配置好的nlp预处理文本
    doc = nlp(text)
    # 传入预处理后的Doc对象,关闭pke内部的normalization
    extractor.load_document(input=doc, language="en", normalization='none')

    # 候选词筛选:保留名词、专有名词、形容词
    pos = {'NOUN', 'PROPN', 'ADJ'}
    extractor.candidate_selection(pos=pos)
    
    # 语法规则筛选(确保规则格式正确)
    extractor.grammar_selection(grammar="NP: {<ADJ>*<NOUN|PROPN>+}")

    # 候选词权重计算,使用TopicRank默认逻辑
    extractor.candidate_weighting(threshold=0.74, method='average')

    # 提取Top10关键词,按需开启冗余移除和词干化
    keyphrases = extractor.get_n_best(n=10, redundancy_removal=True, stemming=True)
    keyphrases = ', '.join([candidate for candidate, weight in keyphrases])
    return keyphrases

关键注意事项

  • 不能排除tagger组件:TopicRank的候选词筛选依赖词性标注,移除后无法识别目标词性,导致提取不到关键词
  • sentencizer是必需的:pke的SpacyDocReader依赖doc.sents拆分句子,必须添加该组件或保留parser
  • max_length需匹配文本长度:确保设置的值大于输入文本的字符数

内容的提问来源于stack exchange,提问作者Mainak Maitra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 19:50:19