处理大量短文本时,序列化spaCy Docs集合的推荐方法是什么?
Great question—this is a common scenario when working with large text collections in spaCy, and there are a few solid approaches to handle it without redundant Vocab/Language saves. The spaCy-native DocBin class is by far the best recommended solution, but I’ll also cover alternatives if you need more flexibility.
Recommended: Use spaCy’s DocBin
DocBin is purpose-built for storing multiple Doc objects efficiently. It automatically shares the Vocab across all Docs, so you only save it once, and it optimizes storage space by avoiding duplicate data. Here’s how to use it:
Saving Docs with DocBin
from spacy.tokens import DocBin import spacy # Assume you have a list of Doc objects (e.g., from nlp.pipe()) nlp = spacy.load("en_core_web_sm") texts = ["This is a sample text", "Another short document", "More text to process"] list_of_docs = list(nlp.pipe(texts)) # Initialize DocBin: specify which annotations to save (optional, saves all if omitted) doc_bin = DocBin(attrs=["LEMMA", "POS", "ENT_IOB", "ENT_TYPE"], docs=list_of_docs) # Save to disk (stores Vocab + all Docs in one file) doc_bin.to_disk("./my_large_doc_collection.spacy")
Loading Docs with DocBin
from spacy.tokens import DocBin import spacy # Load the DocBin from disk doc_bin = DocBin().from_disk("./my_large_doc_collection.spacy") # Get the shared Vocab (if you don't already have it) vocab = doc_bin.vocab # Convert DocBin back to a list of Doc objects list_of_docs = list(doc_bin.get_docs(vocab)) # If you already have an nlp object, you can use its Vocab instead: # nlp = spacy.load("en_core_web_sm") # list_of_docs = list(doc_bin.get_docs(nlp.vocab))
Alternative 1: Save Individual Docs with Shared Vocab
If you prefer to keep each Doc in its own file (useful if you need to access specific Docs without loading the entire collection), you can save the Vocab once and then save each Doc without including the Vocab:
Saving
import spacy from spacy.tokens import Doc nlp = spacy.load("en_core_web_sm") list_of_docs = list(nlp.pipe(["Text 1", "Text 2", "Text 3"])) # Save the Vocab once (critical for loading later) nlp.vocab.to_disk("./shared_vocab") # Save each Doc to a separate file, excluding the Vocab for idx, doc in enumerate(list_of_docs): doc.to_disk(f"./docs/doc_{idx}.spacy", exclude=["vocab"])
Loading
from spacy.tokens import Doc, Vocab # Load the shared Vocab first vocab = Vocab().from_disk("./shared_vocab") # Load each Doc using the shared Vocab loaded_docs = [] for idx in range(3): doc = Doc(vocab).from_disk(f"./docs/doc_{idx}.spacy") loaded_docs.append(doc)
Alternative 2: Serialize Docs to Bytes and Store with Pickle
If you want to store all Docs in a single file without using DocBin, you can serialize each Doc to bytes (excluding the Vocab) and save the list of bytes with pickle:
Saving
import spacy import pickle nlp = spacy.load("en_core_web_sm") list_of_docs = list(nlp.pipe(["Text A", "Text B", "Text C"])) # Save Vocab once nlp.vocab.to_disk("./shared_vocab") # Serialize each Doc to bytes (exclude Vocab to save space) doc_bytes_list = [doc.to_bytes(exclude=["vocab"]) for doc in list_of_docs] # Save the bytes list with pickle with open("./docs_bytes.pkl", "wb") as f: pickle.dump(doc_bytes_list, f)
Loading
from spacy.tokens import Doc, Vocab import pickle # Load Vocab vocab = Vocab().from_disk("./shared_vocab") # Load the bytes list with open("./docs_bytes.pkl", "rb") as f: doc_bytes_list = pickle.load(f) # Reconstruct each Doc loaded_docs = [Doc(vocab).from_bytes(b) for b in doc_bytes_list]
Key Notes
- Always prioritize DocBin: It’s the most efficient, spaCy-native solution, and it handles edge cases (like custom attributes) better than manual serialization.
- Exclude Vocab when saving individual Docs: This cuts down on file size drastically since Vocab is shared across all Docs.
- Custom attributes: If you’ve added custom attributes to your Docs, make sure to register them before loading (using
Doc.set_extension()) to avoid errors.
内容的提问来源于stack exchange,提问作者Bill

