如何使用spaCy实现基于语义而非词汇的CV与招聘启事文本相似度检测?
The core issue you’re facing is that spaCy’s default similarity calculation relies on average word vectors, which ignores word order and context—exactly what the docs warned about. To get similarity based on actual semantic meaning, you need to use context-aware embeddings instead. Here are the most effective solutions:
1. Use spaCy’s Transformer Models (Best Built-in Option)
Transformer-based spaCy models (like en_core_web_trf) generate context-dependent embeddings, meaning the same word gets different vectors based on its surrounding text. This directly addresses your problem by capturing semantic intent rather than just word overlap.
Steps to Implement:
First, install the transformer model:
python -m spacy download en_core_web_trf
Then, load the model and compute similarity as usual—this time, the underlying embeddings will be context-aware:
import spacy # Load transformer-powered pipeline nlp = spacy.load("en_core_web_trf") # Process your CV and job ad texts into Doc objects cv_doc = nlp(pdf_text) ad_doc = nlp(final_text_from_annonce) # Calculate semantic similarity similarity = cv_doc.similarity(ad_doc)
This will produce scores that reflect how well the CV matches the job ad’s actual meaning, not just shared vocabulary. Note: Transformer models are more resource-heavy, but the accuracy gain is worth it for this use case.
2. Use Sentence-BERT for Document-Level Semantic Matching
If you want even better performance for document comparison, consider the sentence-transformers library. It’s optimized specifically for semantic similarity tasks and produces embeddings that capture the overall meaning of a document.
Example Implementation:
Install the library first:
pip install sentence-transformers
Then compute embeddings and similarity:
from sentence_transformers import SentenceTransformer, util # Load a lightweight, pre-trained model tuned for semantic similarity model = SentenceTransformer('all-MiniLM-L6-v2') # Generate embeddings for your CV and job ad cv_embedding = model.encode(pdf_text, convert_to_tensor=True) ad_embedding = model.encode(final_text_from_annonce, convert_to_tensor=True) # Calculate cosine similarity between the embeddings similarity = util.cos_sim(cv_embedding, ad_embedding).item()
This approach is often faster than spaCy’s transformer pipeline and will clearly distinguish between well-matched and poorly-matched CV/ad pairs.
3. (Alternative) Weighted Word Vector Averaging (Less Effective)
If transformers aren’t an option, you can improve the default similarity by weighting word vectors based on their importance (e.g., TF-IDF scores). This gives more weight to rare, meaningful words instead of common stopwords.
Example Code:
import spacy from sklearn.feature_extraction.text import TfidfVectorizer nlp = spacy.load("en_core_web_lg") # Calculate TF-IDF weights for words in your corpus (CV + job ad) corpus = [pdf_text, final_text_from_annonce] tfidf = TfidfVectorizer(tokenizer=lambda text: [token.text for token in nlp(text) if not token.is_stop]) tfidf_matrix = tfidf.fit_transform(corpus) word_weights = dict(zip(tfidf.get_feature_names_out(), tfidf.idf_)) # Function to compute weighted average vector def get_weighted_embedding(text): doc = nlp(text) total_weight = 0.0 weighted_vector = doc.vocab.vectors.zero_vector() for token in doc: if not token.is_stop and token.text in word_weights: weight = word_weights[token.text] weighted_vector += token.vector * weight total_weight += weight return weighted_vector / total_weight if total_weight != 0 else weighted_vector # Calculate similarity between weighted embeddings cv_emb = get_weighted_embedding(pdf_text) ad_emb = get_weighted_embedding(final_text_from_annonce) similarity = cv_emb.dot(ad_emb) / (cv_emb.norm() * ad_emb.norm())
This is a middle-ground solution, but it still doesn’t capture context as effectively as transformers.
Key Takeaway
There’s no "magic parameter" for spaCy’s default similarity() function to switch to semantic matching—you need to use a model that generates context-aware embeddings. The best options are spaCy’s transformer models or Sentence-BERT, both of which will give you the meaningful similarity scores you’re looking for.
内容的提问来源于stack exchange,提问作者Adrien Villemin

