You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何借助Doc2Vec向量计算文档中重要词汇的权重?

Calculating Weights for Key Terms Using Doc2Vec Vectors

Alright, let's break down how to compute meaningful weights for those important terms you've identified with Word2Vec, leveraging vectors from your Doc2Vec model. Here are a few practical, actionable approaches tailored to your existing code setup:


1. Cosine Similarity Between Term Vector and Document Vector

The core intuition here is simple: the more semantically aligned a term's vector is with a document's vector, the more relevant that term is to the document. Doc2Vec's document vectors are trained to encapsulate the overall meaning of the document, so this similarity score makes a natural weight.

Implementation Steps:

First, grab your target document's vector from the model (assuming you stored document tags during training). Then calculate the cosine similarity between the term's inferred vector and the document's vector.

from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

# Load your pre-trained Doc2Vec model (your existing code)
model = Doc2Vec.load(fname)

# Target term and its inferred vector
term = ["suddenly"]
term_vector = model.infer_vector(term).reshape(1, -1)

# Get the document vector (replace "doc_tag_123" with your actual document's training tag)
doc_vector = model.dv["doc_tag_123"].reshape(1, -1)

# Compute the similarity score (this is your term's weight)
term_weight = cosine_similarity(term_vector, doc_vector)[0][0]

print(f"Weight for '{term[0]}': {term_weight:.4f}")

Why This Works:

Cosine similarity ranges from -1 to 1. A score close to 1 means the term and document share strong semantic overlap, so it’s a reliable indicator of the term’s importance to the document.


2. Hybrid Weight: Combine TF with Doc2Vec Similarity

For a more robust weight, you can merge traditional term frequency (TF) with the Doc2Vec similarity score. This accounts for both how often the term appears in the document and how well it matches the document’s core meaning.

Example Code:

from sklearn.feature_extraction.text import TfidfVectorizer

# Assume your target document's text is stored here
document_text = "Your full document content goes here..."

# Calculate raw term frequency (TF) for the term
tf_vectorizer = TfidfVectorizer(use_idf=False, norm=None)
tf_matrix = tf_vectorizer.fit_transform([document_text])
term_tf = tf_matrix[0, tf_vectorizer.vocabulary_.get(term[0], 0)]

# Use the cosine similarity from approach 1
similarity_score = cosine_similarity(term_vector, doc_vector)[0][0]

# Compute hybrid weight: TF * similarity score
hybrid_weight = term_tf * similarity_score

print(f"Hybrid weight for '{term[0]}': {hybrid_weight:.4f}")

Pro Tip: Normalize Weights

If you want to compare weights across multiple terms, normalize the scores to a 0-1 range using min-max scaling—this makes it easier to rank terms by importance.


3. Use Pre-Trained Word Vectors from Doc2Vec

If your Doc2Vec model was trained with the dm=1 (distributed memory) architecture, it learns word embeddings alongside document vectors. You can use these pre-trained word vectors directly instead of inferring a new one, which might be more consistent with the model’s training data.

Code Snippet:

# Check if the term exists in the model's vocabulary
if term[0] in model.wv:
    pre_trained_term_vector = model.wv[term[0]].reshape(1, -1)
    term_weight = cosine_similarity(pre_trained_term_vector, doc_vector)[0][0]
    print(f"Pre-trained vector weight for '{term[0]}': {term_weight:.4f}")
else:
    # Fallback to inferring the vector if the term is out-of-vocabulary
    print(f"Term '{term[0]}' not in vocabulary, using inferred vector as fallback.")

A quick note on infer_vector(): For more accurate results, pass a short context window around the term (instead of just the term itself) if you have access to it. For example: model.infer_vector(["the", "cat", "suddenly", "jumped"]). This helps the model generate a more semantically precise vector.

内容的提问来源于stack exchange,提问作者ucmou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:58:49