使用Spacy v3训练实体链接模型时遭遇ValueError: [E030] 句子边界未设置问题排查
I ran into this exact issue a while back—let's break down why your sentencizer component isn't doing what you expect, and how to fix it.
The Core Problem
Even though you added sentencizer to your pipeline, the gold reference documents used during training (the data from your train/dev corpora) aren't being processed by this component. When the entity linker tries to access eg.reference.sents, those docs don't have is_sent_start markers set, hence the error.
SpaCy's default spacy.Corpus.v1 reader doesn't automatically run your pipeline components on the gold data—it loads the raw annotated docs as-is. Your pipeline components only process the prediction docs during training, not the reference ones.
Solutions
1. Force the Corpus Reader to Process Gold Docs with Sentencizer
Update your [corpora.train] and [corpora.dev] config sections to include an annotate_components setting. This tells spaCy to run the sentencizer (or another component that sets sentence boundaries) on the gold data before training:
[corpora.train] @readers = "spacy.Corpus.v1" path = ${paths.train} max_length = 0 annotate_components = ["sentencizer"] [corpora.dev] @readers = "spacy.Corpus.v1" path = ${paths.dev} max_length = 0 annotate_components = ["sentencizer"]
This will ensure every reference doc has sentence boundaries set when the entity linker tries to access sents.
2. Fix Pipeline Order Conflicts
Your current pipeline order is transformer -> parser -> sentencizer, which is problematic:
- The
parsercomponent already predicts and sets sentence boundaries on its own. - Running
sentencizerafterparsercan overwrite or interfere with the parser's sentence segmentation.
Adjust your pipeline to either:
- Move
sentencizerbeforeparser(so it provides initial sentence splits that the parser refines):[nlp] lang = "en" pipeline = ["transformer","sentencizer","parser","ner", "entity_linker"] batch_size = 128 - Or remove
sentencizerentirely (since the parser handles sentence boundaries):[nlp] lang = "en" pipeline = ["transformer","parser","ner", "entity_linker"] batch_size = 128
If you go with removing sentencizer, update annotate_components to use ["parser"] instead, so the gold docs get sentence boundaries from the parser.
Quick Verification
To confirm your gold data was missing sentence boundaries, run this quick script:
import spacy from spacy.tokens import DocBin # Load a blank nlp with sentencizer nlp = spacy.blank("en") nlp.add_pipe("sentencizer") # Load your training data doc_bin = DocBin().from_disk("path/to/your/train_data.spacy") for doc in doc_bin.get_docs(nlp.vocab): # Check the first 10 tokens' is_sent_start status print([token.is_sent_start for token in doc[:10]])
If you see mostly None values, that confirms the gold docs weren't being processed to set sentence boundaries—fixing the corpus config will resolve this.
内容的提问来源于stack exchange,提问作者Jon Flynn

