关于spaCy中is_oov方法未按预期工作的技术咨询
is_oov Flags Common Words as OOV (v2.0.9) Hey there! Let's break down why your code is returning True for every token, even common words like "dog" or "I".
The Root Cause
The issue comes down to the model you're using. In spaCy 2.x, the en model is an alias for en_core_web_sm—the small English model. This lightweight model doesn't include pre-trained word vectors.
The is_oov attribute doesn't check if a word is a recognized English term (spaCy can still tag/parse those words just fine!). Instead, it checks whether the token has a corresponding entry in the model's pre-trained word vector table. Since the small model has no pre-trained vectors, every token gets marked as out-of-vocabulary.
Fix: Use a Model with Pre-trained Vectors
To get accurate is_oov results, you'll need to use a model that includes pre-trained word vectors (either the medium md or large lg model). Here's how:
Install the compatible model for spaCy 2.0.9:
pip install spacy==2.0.9 python -m spacy download en_core_web_mdLoad the new model in your code:
import spacy nlp = spacy.load('en_core_web_md') doc = nlp('I am sflmgmavknsaccasas dog cat bird bulbasaur') print([tok.is_oov for tok in doc])
Expected Output
Running the updated code will give you:
[False, False, True, False, False, False, True]
- Common words like "I", "am", "dog" return
False(they have pre-trained vectors) - Gibberish ("sflmgmavknsaccasas") and rare terms ("bulbasaur") return
True(no vectors available)
Quick Note
Remember: The small sm model still works great for basic NLP tasks like POS tagging, dependency parsing, and entity recognition. The is_oov flag is only relevant if you're working with word vectors (e.g., similarity checks, vector-based features).
内容的提问来源于stack exchange,提问作者ghonke

