如何在spaCy中处理超长文档?西班牙语词形还原内存溢出问题咨询
西班牙语文本spaCy词形还原长文本内存溢出解决方案
你触发的报错如下:
ValueError: [E088] Text of length 6095095 exceeds maximum of 1000000. The parser and NER models require roughly 1GB of temporary memory per 100,000 characters in the input. This means long texts may cause memory allocation errors. If you're not using the parser or NER, it's probably safe to increase the
nlp.max_lengthlimit. The limit is in number of characters, so you can check whether your inputs are too long by checkinglen(text).
直接调整nlp.max_length但没有关闭高内存占用组件,必然会导致内存耗尽崩溃,可选用以下任意方案解决:
- 分块处理+禁用非必要组件(适配现有spaCy流程)
词形还原不需要依赖parser、NER两个高内存占用组件,加载模型时先禁用这两个组件,再将长文本拆分为单块不超过90万字符的片段逐块处理,最后合并结果即可,无需修改max_length参数:import spacy # 加载模型时禁用parser和NER,内存占用可降低70%以上 nlp = spacy.load("es_core_news_sm", disable=["parser", "ner"]) text = "你的超长西班牙语文本" chunk_size = 900000 all_lemmas = [] # 按固定长度拆分文本,可根据实际情况调整拆分逻辑避免拆断单词 for i in range(0, len(text), chunk_size): chunk = text[i:i+chunk_size] doc = nlp(chunk) all_lemmas.extend([token.lemma_ for token in doc]) - 单独调用spaCy词形还原组件(内存占用最低)
如果仅需要词形还原功能,不需要分词、词性标注等其他能力,可以直接初始化西班牙语的独立词形还原器,完全没有输入长度限制:from spacy.lang.es import Spanish nlp = Spanish() lemmatizer = nlp.add_pipe("lemmatizer") lemmatizer.initialize() # 直接处理文本即可 doc = nlp("你的西班牙语文本") lemmas = [token.lemma_ for token in doc] - 替换轻量级西班牙语文本处理工具
可以换用pattern.es等专门的轻量级西班牙语NLP工具,自带词形还原能力,无输入长度限制,处理速度更快。
内容的提问来源于stack exchange,提问作者Gabriel Costanzo
相关产品推荐
相关产品推荐

