You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Doc2vec对相同文本生成不同向量?如何解决?

Why Identical Texts Get Different Doc2Vec Vectors, and How to Fix It

Hey there, let’s break down why you’re seeing different vectors for identical texts and how to make them match.

The Root Cause

First, let’s look at how your code is set up: you’re creating two TaggedDocument instances with different tags (0 and 1) even though the text is identical.

In Doc2Vec, every unique tag gets its own dedicated document vector that’s trained independently. Even if the input text is the same, the model treats these as separate "documents" (thanks to the distinct tags) and updates each tag’s vector through the training process. While you set a seed, there’s still inherent randomness in initial vector initialization and the order of training updates—this leads to different final vectors for each tag, even with identical input text.

Your code also uses dm_concat=1 (distributed memory mode with concatenation), which combines the document vector with word vectors during training. But since each tag has its own unique vector, the updates for tag 0 and 1 will diverge over epochs, even with the exact same input.

How to Get Identical Vectors for Identical Texts

Here are a few practical approaches to achieve what you want:

1. Assign the Same Tag to Identical Texts

If two texts are identical and you want them to share the same vector, just tag them with the same identifier. Modify your training data creation like this:

from gensim.models.doc2vec import TaggedDocument

f = open('test.txt','r')
trainings = []
for i, data in enumerate(f):
    text = data.strip()
    # Use the text itself as the shared tag for identical content
    trainings.append(TaggedDocument(words=text.split(","), tags=[text]))

model = Doc2Vec(vector_size=5, epochs=55, seed=1, dm_concat=1)
model.build_vocab(trainings)
model.train(trainings, total_examples=model.corpus_count, epochs=model.epochs)

Now both instances of "a" will share the same tag, so the model will train a single vector for that tag—you’ll get the same vector when accessing it.

2. Use infer_vector() for Consistent Text-Based Vectors

If you need to generate vectors for identical texts after training, don’t rely on docvecs (which are tied to training-specific tags). Instead, use the infer_vector() method, which generates a vector based purely on the text content. For identical texts, this will produce nearly identical vectors (you can eliminate randomness by setting a seed and increasing inference steps):

model = Doc2Vec.load('doc2vec.model')
# Infer vector for the text "a" with fixed seed and sufficient steps
vec1 = model.infer_vector(["a"], steps=100, seed=1)
vec2 = model.infer_vector(["a"], steps=100, seed=1)
print(vec1)
print(vec2)  # These will be identical (or extremely close)

3. Average Vectors of Identical Texts (If You Can’t Retrain)

If you already have a trained model with different tags for identical texts, you can compute the average of all vectors linked to the same text to get a single representative vector:

from collections import defaultdict

text_to_vecs = defaultdict(list)
f = open('test.txt','r')
for i, data in enumerate(f):
    text = data.strip()
    text_to_vecs[text].append(model.docvecs[i])

# Calculate average vector for each unique text
for text, vecs in text_to_vecs.items():
    avg_vec = sum(vecs) / len(vecs)
    print(f"Text: {text}, Average Vector: {avg_vec}")

Quick Note on Doc2Vec Behavior

Remember: Doc2Vec’s docvecs are designed to represent individual documents (each with a unique tag), not unique text content. If your goal is to get vectors for distinct text strings rather than individual document entries, using infer_vector() or shared tags is the right approach.

内容的提问来源于stack exchange,提问作者Thanh Bui

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 03:54:11