如何借助Doc2Vec向量计算文档中重要词汇的权重?
Alright, let's break down how to compute meaningful weights for those important terms you've identified with Word2Vec, leveraging vectors from your Doc2Vec model. Here are a few practical, actionable approaches tailored to your existing code setup:
1. Cosine Similarity Between Term Vector and Document Vector
The core intuition here is simple: the more semantically aligned a term's vector is with a document's vector, the more relevant that term is to the document. Doc2Vec's document vectors are trained to encapsulate the overall meaning of the document, so this similarity score makes a natural weight.
Implementation Steps:
First, grab your target document's vector from the model (assuming you stored document tags during training). Then calculate the cosine similarity between the term's inferred vector and the document's vector.
from sklearn.metrics.pairwise import cosine_similarity import numpy as np # Load your pre-trained Doc2Vec model (your existing code) model = Doc2Vec.load(fname) # Target term and its inferred vector term = ["suddenly"] term_vector = model.infer_vector(term).reshape(1, -1) # Get the document vector (replace "doc_tag_123" with your actual document's training tag) doc_vector = model.dv["doc_tag_123"].reshape(1, -1) # Compute the similarity score (this is your term's weight) term_weight = cosine_similarity(term_vector, doc_vector)[0][0] print(f"Weight for '{term[0]}': {term_weight:.4f}")
Why This Works:
Cosine similarity ranges from -1 to 1. A score close to 1 means the term and document share strong semantic overlap, so it’s a reliable indicator of the term’s importance to the document.
2. Hybrid Weight: Combine TF with Doc2Vec Similarity
For a more robust weight, you can merge traditional term frequency (TF) with the Doc2Vec similarity score. This accounts for both how often the term appears in the document and how well it matches the document’s core meaning.
Example Code:
from sklearn.feature_extraction.text import TfidfVectorizer # Assume your target document's text is stored here document_text = "Your full document content goes here..." # Calculate raw term frequency (TF) for the term tf_vectorizer = TfidfVectorizer(use_idf=False, norm=None) tf_matrix = tf_vectorizer.fit_transform([document_text]) term_tf = tf_matrix[0, tf_vectorizer.vocabulary_.get(term[0], 0)] # Use the cosine similarity from approach 1 similarity_score = cosine_similarity(term_vector, doc_vector)[0][0] # Compute hybrid weight: TF * similarity score hybrid_weight = term_tf * similarity_score print(f"Hybrid weight for '{term[0]}': {hybrid_weight:.4f}")
Pro Tip: Normalize Weights
If you want to compare weights across multiple terms, normalize the scores to a 0-1 range using min-max scaling—this makes it easier to rank terms by importance.
3. Use Pre-Trained Word Vectors from Doc2Vec
If your Doc2Vec model was trained with the dm=1 (distributed memory) architecture, it learns word embeddings alongside document vectors. You can use these pre-trained word vectors directly instead of inferring a new one, which might be more consistent with the model’s training data.
Code Snippet:
# Check if the term exists in the model's vocabulary if term[0] in model.wv: pre_trained_term_vector = model.wv[term[0]].reshape(1, -1) term_weight = cosine_similarity(pre_trained_term_vector, doc_vector)[0][0] print(f"Pre-trained vector weight for '{term[0]}': {term_weight:.4f}") else: # Fallback to inferring the vector if the term is out-of-vocabulary print(f"Term '{term[0]}' not in vocabulary, using inferred vector as fallback.")
A quick note on infer_vector(): For more accurate results, pass a short context window around the term (instead of just the term itself) if you have access to it. For example: model.infer_vector(["the", "cat", "suddenly", "jumped"]). This helps the model generate a more semantically precise vector.
内容的提问来源于stack exchange,提问作者ucmou

