You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

处理大量短文本时,序列化spaCy Docs集合的推荐方法是什么?

Efficiently Saving and Loading Multiple spaCy Doc Objects

Great question—this is a common scenario when working with large text collections in spaCy, and there are a few solid approaches to handle it without redundant Vocab/Language saves. The spaCy-native DocBin class is by far the best recommended solution, but I’ll also cover alternatives if you need more flexibility.

DocBin is purpose-built for storing multiple Doc objects efficiently. It automatically shares the Vocab across all Docs, so you only save it once, and it optimizes storage space by avoiding duplicate data. Here’s how to use it:

Saving Docs with DocBin

from spacy.tokens import DocBin
import spacy

# Assume you have a list of Doc objects (e.g., from nlp.pipe())
nlp = spacy.load("en_core_web_sm")
texts = ["This is a sample text", "Another short document", "More text to process"]
list_of_docs = list(nlp.pipe(texts))

# Initialize DocBin: specify which annotations to save (optional, saves all if omitted)
doc_bin = DocBin(attrs=["LEMMA", "POS", "ENT_IOB", "ENT_TYPE"], docs=list_of_docs)

# Save to disk (stores Vocab + all Docs in one file)
doc_bin.to_disk("./my_large_doc_collection.spacy")

Loading Docs with DocBin

from spacy.tokens import DocBin
import spacy

# Load the DocBin from disk
doc_bin = DocBin().from_disk("./my_large_doc_collection.spacy")

# Get the shared Vocab (if you don't already have it)
vocab = doc_bin.vocab

# Convert DocBin back to a list of Doc objects
list_of_docs = list(doc_bin.get_docs(vocab))

# If you already have an nlp object, you can use its Vocab instead:
# nlp = spacy.load("en_core_web_sm")
# list_of_docs = list(doc_bin.get_docs(nlp.vocab))

Alternative 1: Save Individual Docs with Shared Vocab

If you prefer to keep each Doc in its own file (useful if you need to access specific Docs without loading the entire collection), you can save the Vocab once and then save each Doc without including the Vocab:

Saving

import spacy
from spacy.tokens import Doc

nlp = spacy.load("en_core_web_sm")
list_of_docs = list(nlp.pipe(["Text 1", "Text 2", "Text 3"]))

# Save the Vocab once (critical for loading later)
nlp.vocab.to_disk("./shared_vocab")

# Save each Doc to a separate file, excluding the Vocab
for idx, doc in enumerate(list_of_docs):
    doc.to_disk(f"./docs/doc_{idx}.spacy", exclude=["vocab"])

Loading

from spacy.tokens import Doc, Vocab

# Load the shared Vocab first
vocab = Vocab().from_disk("./shared_vocab")

# Load each Doc using the shared Vocab
loaded_docs = []
for idx in range(3):
    doc = Doc(vocab).from_disk(f"./docs/doc_{idx}.spacy")
    loaded_docs.append(doc)

Alternative 2: Serialize Docs to Bytes and Store with Pickle

If you want to store all Docs in a single file without using DocBin, you can serialize each Doc to bytes (excluding the Vocab) and save the list of bytes with pickle:

Saving

import spacy
import pickle

nlp = spacy.load("en_core_web_sm")
list_of_docs = list(nlp.pipe(["Text A", "Text B", "Text C"]))

# Save Vocab once
nlp.vocab.to_disk("./shared_vocab")

# Serialize each Doc to bytes (exclude Vocab to save space)
doc_bytes_list = [doc.to_bytes(exclude=["vocab"]) for doc in list_of_docs]

# Save the bytes list with pickle
with open("./docs_bytes.pkl", "wb") as f:
    pickle.dump(doc_bytes_list, f)

Loading

from spacy.tokens import Doc, Vocab
import pickle

# Load Vocab
vocab = Vocab().from_disk("./shared_vocab")

# Load the bytes list
with open("./docs_bytes.pkl", "rb") as f:
    doc_bytes_list = pickle.load(f)

# Reconstruct each Doc
loaded_docs = [Doc(vocab).from_bytes(b) for b in doc_bytes_list]

Key Notes

  • Always prioritize DocBin: It’s the most efficient, spaCy-native solution, and it handles edge cases (like custom attributes) better than manual serialization.
  • Exclude Vocab when saving individual Docs: This cuts down on file size drastically since Vocab is shared across all Docs.
  • Custom attributes: If you’ve added custom attributes to your Docs, make sure to register them before loading (using Doc.set_extension()) to avoid errors.

内容的提问来源于stack exchange,提问作者Bill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:43:56