NLP入门:如何用Word2Vec计算句子相似度及实现Gensim Word2Vec模型
Hey there! Let's break down your questions and fix up your code to meet all three of your goals. As an NLP beginner, you're asking three great foundational questions—let's tackle them one by one, starting with the most critical part: getting a working Word2Vec model set up.
1. Implementing a Gensim Word2Vec Model
First off, your current code has a key issue: you're trying to call word2vec(w) directly, but word2vec is the Gensim module, not a function to fetch word vectors. You need to either train your own Word2Vec model on text data or load a pre-trained one. Let's start with training a small model for learning purposes.
Training a Basic Word2Vec Model
To train a model, you need a corpus (a collection of sentences). Here's how to set that up:
from gensim.models import Word2Vec import numpy as np from sklearn.metrics.pairwise import cosine_similarity # Sample corpus (expand this with more text for better vector quality!) corpus = [ "I am going to India", "I am going to Bharat", "India is a country in South Asia", "Bharat is the Hindi name for India", "People travel to India for holidays" ] # Preprocess: split each sentence into lowercase words (matches model training requirements) tokenized_corpus = [sentence.lower().split() for sentence in corpus] # Train the Word2Vec model model = Word2Vec( sentences=tokenized_corpus, vector_size=100, # Length of each word vector window=5, # Number of surrounding words to consider for context min_count=1, # Ignore words that appear fewer than this many times workers=4 # Number of threads to speed up training ) # Optional: Save the model for later use # model.save("my_word2vec_model.model") # Optional: Load a saved model # model = Word2Vec.load("my_word2vec_model.model")
2. Printing Word-Level Similarity Scores
Now that we have a trained model, we can calculate similarity between individual words. For your two sentences, we can compare corresponding word pairs or find the most similar words across sentences.
Comparing Corresponding Word Pairs
Here's how to loop through your sentences and print similarity scores for each matching word pair:
sentence1 = "I am going to India" sentence2 = "I am going to Bharat" # Tokenize and lowercase to match model preprocessing words1 = sentence1.lower().split() words2 = sentence2.lower().split() print("=== Word-Level Similarity Scores ===") for w1, w2 in zip(words1, words2): try: # Fetch similarity score between the two words sim_score = model.wv.similarity(w1, w2) print(f"Similarity between '{w1}' and '{w2}': {sim_score:.4f}") except KeyError as e: print(f"Warning: Word '{e}' isn't in the model's vocabulary")
Finding Most Similar Words (for mismatched sentence lengths)
If your sentences don't have the same number of words, you can find the closest match in the second sentence for each word in the first:
print("\n=== Most Similar Words Across Sentences ===") for w1 in words1: if w1 in model.wv: # Get the top 1 most similar word from the model's vocabulary most_similar = model.wv.most_similar(w1, topn=1)[0] print(f"'{w1}' is most similar to '{most_similar[0]}' with score: {most_similar[1]:.4f}") else: print(f"Word '{w1}' not found in vocabulary")
3. Calculating Sentence Similarity
Your original idea of using the average of word vectors is a solid baseline for sentence similarity. Let's fix your code to use the trained model's vectors correctly.
Step 1: Create Sentence Vectors (Average of Word Vectors)
def get_sentence_vector(sentence, model): words = sentence.lower().split() # Collect vectors only for words that exist in the model word_vectors = [model.wv[word] for word in words if word in model.wv] if not word_vectors: # Return a zero vector if no words are in the vocabulary return np.zeros(model.vector_size) # Calculate the average vector for the sentence return np.mean(word_vectors, axis=0) # Generate vectors for your two sentences sentence1_vec = get_sentence_vector(sentence1, model) sentence2_vec = get_sentence_vector(sentence2, model)
Step 2: Calculate Cosine Similarity
You can implement this manually (like your original attempt) or use sklearn's built-in function for simplicity:
# Manual cosine similarity calculation def cosine_similarity_manual(vec1, vec2): dot_product = np.dot(vec1, vec2) norm1 = np.linalg.norm(vec1) norm2 = np.linalg.norm(vec2) if norm1 == 0 or norm2 == 0: return 0.0 return dot_product / (norm1 * norm2) # Calculate and print results manual_sim = cosine_similarity_manual(sentence1_vec, sentence2_vec) sklearn_sim = cosine_similarity([sentence1_vec], [sentence2_vec])[0][0] print("\n=== Sentence Similarity ===") print(f"Manual calculation: {manual_sim:.4f}") print(f"Sklearn calculation: {sklearn_sim:.4f}")
Full Working Code
Putting it all together, here's the complete script:
from gensim.models import Word2Vec import numpy as np from sklearn.metrics.pairwise import cosine_similarity # Train Word2Vec model corpus = [ "I am going to India", "I am going to Bharat", "India is a country in South Asia", "Bharat is the Hindi name for India", "People travel to India for holidays" ] tokenized_corpus = [sent.lower().split() for sent in corpus] model = Word2Vec(tokenized_corpus, vector_size=100, window=5, min_count=1, workers=4) # Define target sentences sentence1 = "I am going to India" sentence2 = "I am going to Bharat" words1 = sentence1.lower().split() words2 = sentence2.lower().split() # Print word-level similarities print("=== Word-Level Similarity Scores ===") for w1, w2 in zip(words1, words2): if w1 in model.wv and w2 in model.wv: sim = model.wv.similarity(w1, w2) print(f"'{w1}' vs '{w2}': {sim:.4f}") else: missing = [w for w in [w1, w2] if w not in model.wv] print(f"Missing words in vocabulary: {', '.join(missing)}") # Calculate sentence similarity def get_sentence_vec(sentence, model): vecs = [model.wv[w] for w in sentence.lower().split() if w in model.wv] return np.mean(vecs, axis=0) if vecs else np.zeros(model.vector_size) vec1 = get_sentence_vec(sentence1, model) vec2 = get_sentence_vec(sentence2, model) sentence_sim = cosine_similarity([vec1], [vec2])[0][0] print(f"\n=== Sentence Similarity ===") print(f"Similarity between the two sentences: {sentence_sim:.4f}")
Quick Tips for You
- Vocabulary Coverage: If a word isn't in your model's vocabulary, you can't get its vector. Training on more diverse text will improve this.
- Preprocessing Consistency: Always align your input text with how you preprocessed the training corpus (lowercasing, tokenization, etc.).
- Advanced Sentence Similarity: For better accuracy later, explore methods like Doc2Vec, BERT embeddings, or Sentence-BERT—but Word2Vec averages are a fantastic starting point.
内容的提问来源于stack exchange,提问作者marton mar suri

