You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gensim Doc2Vec能否区分核心相同但正负语境不同的语句?

Can Gensim's Doc2Vec Distinguish Sentences with Opposite Context but Same Core Content?

Great question! Let's break down how Doc2Vec behaves in this scenario, using your example sentences:

"I love Machine Learning" (Sentence A)
"I do not love Machine Learning" (Sentence B)

Short Answer

Doc2Vec can distinguish these two sentences under the right conditions, but it won’t necessarily give a negative cosine similarity. Instead, the similarity score will be noticeably lower than 1 (the score for identical sentences), reflecting the semantic difference introduced by the negation.

Detailed Explanation

Doc2Vec learns document embeddings by treating each document as a unique "token" alongside regular words during training. Its ability to pick up on negation depends on a few key factors:

  • Training Data & Context: If your training corpus includes plenty of sentences with negation structures (like "X do not Y" vs "X Y"), the model will learn to associate negation words like "not" with reversed semantic meaning. Without this context, it might prioritize shared core words ("love", "Machine Learning") and treat the sentences as more similar than they are.
  • Model Architecture: The dm=1 (Distributed Memory) mode tends to perform better on capturing word order and contextual nuances like negation, compared to dbow=1 (Distributed Bag of Words) which focuses more on word frequency.
  • Training Sufficiently: With enough epochs and a reasonable vector size, the model has more capacity to fine-tune the embeddings to reflect subtle semantic differences.

Example Code & Expected Outcome

Here’s a quick snippet to test this with your sentences:

from gensim.models.doc2vec import Doc2Vec, TaggedDocument
from nltk.tokenize import word_tokenize
from sklearn.metrics.pairwise import cosine_similarity

# Prepare tagged documents
docs = [
    TaggedDocument(word_tokenize("I love Machine Learning"), tags=["doc_a"]),
    TaggedDocument(word_tokenize("I do not love Machine Learning"), tags=["doc_b"])
]

# Train a small Doc2Vec model
model = Doc2Vec(vector_size=50, min_count=1, epochs=100, dm=1)
model.build_vocab(docs)
model.train(docs, total_examples=model.corpus_count, epochs=model.epochs)

# Get embeddings and calculate similarity
vec_a = model.dv["doc_a"].reshape(1, -1)
vec_b = model.dv["doc_b"].reshape(1, -1)
similarity_score = cosine_similarity(vec_a, vec_b)[0][0]

print(f"Cosine Similarity between A and B: {round(similarity_score, 2)}")

When you run this, you’ll likely get a score around 0.4–0.6 (not close to 1, but also not negative). This shows the model recognizes the negation creates a meaningful semantic gap, but since most words are shared, the vectors aren’t fully opposite (which would require a score near -1).

Key Takeaway

Doc2Vec doesn’t have built-in negation handling, but it can learn to distinguish negated vs non-negated sentences if trained on appropriate data. The similarity score won’t be negative (that’s reserved for truly opposing concepts), but it will be low enough to tell the two sentences apart.

内容的提问来源于stack exchange,提问作者DK818

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:40:29