Gensim Doc2Vec能否区分核心相同但正负语境不同的语句?
Great question! Let's break down how Doc2Vec behaves in this scenario, using your example sentences:
"I love Machine Learning" (Sentence A)
"I do not love Machine Learning" (Sentence B)
Short Answer
Doc2Vec can distinguish these two sentences under the right conditions, but it won’t necessarily give a negative cosine similarity. Instead, the similarity score will be noticeably lower than 1 (the score for identical sentences), reflecting the semantic difference introduced by the negation.
Detailed Explanation
Doc2Vec learns document embeddings by treating each document as a unique "token" alongside regular words during training. Its ability to pick up on negation depends on a few key factors:
- Training Data & Context: If your training corpus includes plenty of sentences with negation structures (like "X do not Y" vs "X Y"), the model will learn to associate negation words like "not" with reversed semantic meaning. Without this context, it might prioritize shared core words ("love", "Machine Learning") and treat the sentences as more similar than they are.
- Model Architecture: The
dm=1(Distributed Memory) mode tends to perform better on capturing word order and contextual nuances like negation, compared todbow=1(Distributed Bag of Words) which focuses more on word frequency. - Training Sufficiently: With enough epochs and a reasonable vector size, the model has more capacity to fine-tune the embeddings to reflect subtle semantic differences.
Example Code & Expected Outcome
Here’s a quick snippet to test this with your sentences:
from gensim.models.doc2vec import Doc2Vec, TaggedDocument from nltk.tokenize import word_tokenize from sklearn.metrics.pairwise import cosine_similarity # Prepare tagged documents docs = [ TaggedDocument(word_tokenize("I love Machine Learning"), tags=["doc_a"]), TaggedDocument(word_tokenize("I do not love Machine Learning"), tags=["doc_b"]) ] # Train a small Doc2Vec model model = Doc2Vec(vector_size=50, min_count=1, epochs=100, dm=1) model.build_vocab(docs) model.train(docs, total_examples=model.corpus_count, epochs=model.epochs) # Get embeddings and calculate similarity vec_a = model.dv["doc_a"].reshape(1, -1) vec_b = model.dv["doc_b"].reshape(1, -1) similarity_score = cosine_similarity(vec_a, vec_b)[0][0] print(f"Cosine Similarity between A and B: {round(similarity_score, 2)}")
When you run this, you’ll likely get a score around 0.4–0.6 (not close to 1, but also not negative). This shows the model recognizes the negation creates a meaningful semantic gap, but since most words are shared, the vectors aren’t fully opposite (which would require a score near -1).
Key Takeaway
Doc2Vec doesn’t have built-in negation handling, but it can learn to distinguish negated vs non-negated sentences if trained on appropriate data. The similarity score won’t be negative (that’s reserved for truly opposing concepts), but it will be low enough to tell the two sentences apart.
内容的提问来源于stack exchange,提问作者DK818

