基于Word2Vec的句子相似度计算:代码运行失败求助
Hey there, fellow researcher! Let's troubleshoot that Word2Vec sentence similarity code you're working on. It looks like you've got a solid start, but there are a few common pitfalls that might be causing it to fail. Let's walk through fixing it step by step.
First, let's address the incomplete part in your code snippet and fill in the gaps, plus fix the common issues that break execution:
1. Complete & Correct the Feature Vector Calculation
Your code cuts off mid-check for word membership in the vocabulary set. Here's the full, error-fixed version of the avg_feature_vector function:
import numpy as np from scipy import spatial def avg_feature_vector(sentence, model, num_features, index2word_set): words = sentence.split() feature_vec = np.zeros((num_features, ), dtype='float32') n_words = 0 for word in words: # Fix: Complete the vocabulary check (you had a typo/cut-off here) if word in index2word_set: feature_vec += model.wv[word] n_words += 1 # Handle edge case: no words from the sentence exist in the model's vocabulary if n_words > 0: feature_vec = feature_vec / n_words return feature_vec
2. Critical Pre-Requisites You Might Have Missed
Before running the function, make sure you've covered these often-overlooked steps:
- Load/Train a Valid Word2Vec Model: You need either a custom-trained model on your research corpus or a pre-trained one (like Google's 300-dimensional Word2Vec model). Example code for loading:
from gensim.models import Word2Vec from gensim.models.keyedvectors import KeyedVectors # For a custom-trained model: # model = Word2Vec.load("your_trained_model_path.model") # For a pre-trained binary model (e.g., Google's): # model = KeyedVectors.load_word2vec_format('GoogleNews-vectors-negative300.bin', binary=True) - Fix Vocabulary Set Initialization: Newer Gensim versions renamed
index2wordtoindex_to_key— use the correct attribute to avoid AttributeErrors:index2word_set = set(model.wv.index_to_key) # Use model.wv.index2word for older Gensim versions - Match Feature Dimension: The
num_featuresparameter must exactly match the dimensionality of your Word2Vec model (e.g., 300 for Google's pre-trained model).
3. Calculate Sentence Similarity
Once you have valid feature vectors for your sentences, use cosine similarity to compare them:
# Example usage with sample sentences sentence1 = "Your first research sentence" sentence2 = "Your second comparison sentence" num_features = 300 # Match your model's dimension # Generate average feature vectors for both sentences vec1 = avg_feature_vector(sentence1, model, num_features, index2word_set) vec2 = avg_feature_vector(sentence2, model, num_features, index2word_set) # Compute cosine similarity (1 - distance = similarity score) similarity_score = 1 - spatial.distance.cosine(vec1, vec2) print(f"Sentence Similarity Score: {similarity_score:.4f}")
4. Common Error Fixes to Keep in Mind
- KeyError for Unseen Words: The function skips words not in the model's vocabulary (which is intentional), but if a sentence has no overlapping words, it returns a zero vector. Add a warning for this edge case if needed.
- Gensim Version Conflicts: If you get an AttributeError for
index2word, switch toindex_to_key(this is a breaking change in Gensim 4.0+). - Data Type Mismatches: Using
float32for the feature vector ensures compatibility with most Word2Vec model vector dtypes.
5. Bonus: Research-Grade Improvements
Since you're a grad student, here are tweaks to boost performance for your work:
- Text Preprocessing: Add lowercasing, stopword removal, or lemmatization (using NLTK or spaCy) before splitting sentences to reduce noise.
- Weighted Averages: Replace simple averaging with TF-IDF-weighted averages to prioritize semantically important words.
- Alternative Approaches: If you can expand beyond basic Word2Vec, consider Doc2Vec (built for sentence-level embeddings) or Sentence-BERT for better accuracy — but the above method works perfectly if you need to stick to Word2Vec.
内容的提问来源于stack exchange,提问作者pintu lohar

