You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用spaCy实现基于语义而非词汇的CV与招聘启事文本相似度检测?

How to Get Semantic Similarity (Not Just Word Overlap) with spaCy

The core issue you’re facing is that spaCy’s default similarity calculation relies on average word vectors, which ignores word order and context—exactly what the docs warned about. To get similarity based on actual semantic meaning, you need to use context-aware embeddings instead. Here are the most effective solutions:

1. Use spaCy’s Transformer Models (Best Built-in Option)

Transformer-based spaCy models (like en_core_web_trf) generate context-dependent embeddings, meaning the same word gets different vectors based on its surrounding text. This directly addresses your problem by capturing semantic intent rather than just word overlap.

Steps to Implement:

First, install the transformer model:

python -m spacy download en_core_web_trf

Then, load the model and compute similarity as usual—this time, the underlying embeddings will be context-aware:

import spacy

# Load transformer-powered pipeline
nlp = spacy.load("en_core_web_trf")

# Process your CV and job ad texts into Doc objects
cv_doc = nlp(pdf_text)
ad_doc = nlp(final_text_from_annonce)

# Calculate semantic similarity
similarity = cv_doc.similarity(ad_doc)

This will produce scores that reflect how well the CV matches the job ad’s actual meaning, not just shared vocabulary. Note: Transformer models are more resource-heavy, but the accuracy gain is worth it for this use case.

2. Use Sentence-BERT for Document-Level Semantic Matching

If you want even better performance for document comparison, consider the sentence-transformers library. It’s optimized specifically for semantic similarity tasks and produces embeddings that capture the overall meaning of a document.

Example Implementation:

Install the library first:

pip install sentence-transformers

Then compute embeddings and similarity:

from sentence_transformers import SentenceTransformer, util

# Load a lightweight, pre-trained model tuned for semantic similarity
model = SentenceTransformer('all-MiniLM-L6-v2')

# Generate embeddings for your CV and job ad
cv_embedding = model.encode(pdf_text, convert_to_tensor=True)
ad_embedding = model.encode(final_text_from_annonce, convert_to_tensor=True)

# Calculate cosine similarity between the embeddings
similarity = util.cos_sim(cv_embedding, ad_embedding).item()

This approach is often faster than spaCy’s transformer pipeline and will clearly distinguish between well-matched and poorly-matched CV/ad pairs.

3. (Alternative) Weighted Word Vector Averaging (Less Effective)

If transformers aren’t an option, you can improve the default similarity by weighting word vectors based on their importance (e.g., TF-IDF scores). This gives more weight to rare, meaningful words instead of common stopwords.

Example Code:

import spacy
from sklearn.feature_extraction.text import TfidfVectorizer

nlp = spacy.load("en_core_web_lg")

# Calculate TF-IDF weights for words in your corpus (CV + job ad)
corpus = [pdf_text, final_text_from_annonce]
tfidf = TfidfVectorizer(tokenizer=lambda text: [token.text for token in nlp(text) if not token.is_stop])
tfidf_matrix = tfidf.fit_transform(corpus)
word_weights = dict(zip(tfidf.get_feature_names_out(), tfidf.idf_))

# Function to compute weighted average vector
def get_weighted_embedding(text):
    doc = nlp(text)
    total_weight = 0.0
    weighted_vector = doc.vocab.vectors.zero_vector()
    for token in doc:
        if not token.is_stop and token.text in word_weights:
            weight = word_weights[token.text]
            weighted_vector += token.vector * weight
            total_weight += weight
    return weighted_vector / total_weight if total_weight != 0 else weighted_vector

# Calculate similarity between weighted embeddings
cv_emb = get_weighted_embedding(pdf_text)
ad_emb = get_weighted_embedding(final_text_from_annonce)
similarity = cv_emb.dot(ad_emb) / (cv_emb.norm() * ad_emb.norm())

This is a middle-ground solution, but it still doesn’t capture context as effectively as transformers.

Key Takeaway

There’s no "magic parameter" for spaCy’s default similarity() function to switch to semantic matching—you need to use a model that generates context-aware embeddings. The best options are spaCy’s transformer models or Sentence-BERT, both of which will give you the meaningful similarity scores you’re looking for.

内容的提问来源于stack exchange,提问作者Adrien Villemin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 16:10:28