You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在spaCy中处理大型数据集?解决大文件内存问题

解决spaCy处理大文本的内存问题

针对你遇到的大文件处理内存溢出问题,有几个实用的解决办法:

1. 逐行处理CSV文件

不要一次性读取整个50MB的CSV文件,而是逐行读取并处理,这样每次只在内存中保留一行数据,内存占用极低。修改后的代码如下:

import re
import spacy

nlp = spacy.load("de_core_news_sm")

with open(".data.csv", "r", encoding="utf-8") as file:
    for line in file:
        # 清理当前行文本
        cleaned_line = re.sub(r"[^a-zA-Z0-9ß\.,!\?-]", " ", line)
        cleaned_line = cleaned_line.lower()
        # 处理并打印分词
        doc = nlp(cleaned_line)
        for token in doc:
            print(token.text)

2. 禁用spaCy不需要的管道

你只需要分词功能,不需要parser(句法分析)和NER(命名实体识别),这些模块会占用大量内存。加载模型时直接禁用它们,只保留tokenizer:

import re
import spacy

# 禁用不需要的管道,只保留分词器
nlp = spacy.load("de_core_news_sm", disable=["parser", "ner"])

with open(".data.csv", "r", encoding="utf-8") as file:
    text = file.read()
text = re.sub(r"[^a-zA-Z0-9ß\.,!\?-]", " ", text)
text = text.lower()
doc = nlp(text)
for token in doc:
    print(token.text)

这种方法即使一次性读取文件,内存占用也会大幅降低,因为不需要加载和运行那些重型模块。

3. 分块处理大文本

如果必须处理整块文本,把大文本分割成多个小的块(比如每100000字符为一块),逐个处理:

import re
import spacy

nlp = spacy.load("de_core_news_sm", disable=["parser", "ner"])
chunk_size = 100000  # 定义每个块的字符数

with open(".data.csv", "r", encoding="utf-8") as file:
    text = file.read()
text = re.sub(r"[^a-zA-Z0-9ß\.,!\?-]", " ", text)
text = text.lower()

# 分割文本为多个块
for i in range(0, len(text), chunk_size):
    chunk = text[i:i+chunk_size]
    doc = nlp(chunk)
    for token in doc:
        print(token.text)

内容的提问来源于stack exchange,提问作者JukeboxHero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 18:25:36