为何Doc2vec对相同文本生成不同向量?如何解决?
Hey there, let’s break down why you’re seeing different vectors for identical texts and how to make them match.
The Root Cause
First, let’s look at how your code is set up: you’re creating two TaggedDocument instances with different tags (0 and 1) even though the text is identical.
In Doc2Vec, every unique tag gets its own dedicated document vector that’s trained independently. Even if the input text is the same, the model treats these as separate "documents" (thanks to the distinct tags) and updates each tag’s vector through the training process. While you set a seed, there’s still inherent randomness in initial vector initialization and the order of training updates—this leads to different final vectors for each tag, even with identical input text.
Your code also uses dm_concat=1 (distributed memory mode with concatenation), which combines the document vector with word vectors during training. But since each tag has its own unique vector, the updates for tag 0 and 1 will diverge over epochs, even with the exact same input.
How to Get Identical Vectors for Identical Texts
Here are a few practical approaches to achieve what you want:
1. Assign the Same Tag to Identical Texts
If two texts are identical and you want them to share the same vector, just tag them with the same identifier. Modify your training data creation like this:
from gensim.models.doc2vec import TaggedDocument f = open('test.txt','r') trainings = [] for i, data in enumerate(f): text = data.strip() # Use the text itself as the shared tag for identical content trainings.append(TaggedDocument(words=text.split(","), tags=[text])) model = Doc2Vec(vector_size=5, epochs=55, seed=1, dm_concat=1) model.build_vocab(trainings) model.train(trainings, total_examples=model.corpus_count, epochs=model.epochs)
Now both instances of "a" will share the same tag, so the model will train a single vector for that tag—you’ll get the same vector when accessing it.
2. Use infer_vector() for Consistent Text-Based Vectors
If you need to generate vectors for identical texts after training, don’t rely on docvecs (which are tied to training-specific tags). Instead, use the infer_vector() method, which generates a vector based purely on the text content. For identical texts, this will produce nearly identical vectors (you can eliminate randomness by setting a seed and increasing inference steps):
model = Doc2Vec.load('doc2vec.model') # Infer vector for the text "a" with fixed seed and sufficient steps vec1 = model.infer_vector(["a"], steps=100, seed=1) vec2 = model.infer_vector(["a"], steps=100, seed=1) print(vec1) print(vec2) # These will be identical (or extremely close)
3. Average Vectors of Identical Texts (If You Can’t Retrain)
If you already have a trained model with different tags for identical texts, you can compute the average of all vectors linked to the same text to get a single representative vector:
from collections import defaultdict text_to_vecs = defaultdict(list) f = open('test.txt','r') for i, data in enumerate(f): text = data.strip() text_to_vecs[text].append(model.docvecs[i]) # Calculate average vector for each unique text for text, vecs in text_to_vecs.items(): avg_vec = sum(vecs) / len(vecs) print(f"Text: {text}, Average Vector: {avg_vec}")
Quick Note on Doc2Vec Behavior
Remember: Doc2Vec’s docvecs are designed to represent individual documents (each with a unique tag), not unique text content. If your goal is to get vectors for distinct text strings rather than individual document entries, using infer_vector() or shared tags is the right approach.
内容的提问来源于stack exchange,提问作者Thanh Bui

