You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP入门:如何用Word2Vec计算句子相似度及实现Gensim Word2Vec模型

Hey there! Let's break down your questions and fix up your code to meet all three of your goals. As an NLP beginner, you're asking three great foundational questions—let's tackle them one by one, starting with the most critical part: getting a working Word2Vec model set up.

1. Implementing a Gensim Word2Vec Model

First off, your current code has a key issue: you're trying to call word2vec(w) directly, but word2vec is the Gensim module, not a function to fetch word vectors. You need to either train your own Word2Vec model on text data or load a pre-trained one. Let's start with training a small model for learning purposes.

Training a Basic Word2Vec Model

To train a model, you need a corpus (a collection of sentences). Here's how to set that up:

from gensim.models import Word2Vec
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity

# Sample corpus (expand this with more text for better vector quality!)
corpus = [
    "I am going to India",
    "I am going to Bharat",
    "India is a country in South Asia",
    "Bharat is the Hindi name for India",
    "People travel to India for holidays"
]

# Preprocess: split each sentence into lowercase words (matches model training requirements)
tokenized_corpus = [sentence.lower().split() for sentence in corpus]

# Train the Word2Vec model
model = Word2Vec(
    sentences=tokenized_corpus,
    vector_size=100,  # Length of each word vector
    window=5,         # Number of surrounding words to consider for context
    min_count=1,      # Ignore words that appear fewer than this many times
    workers=4         # Number of threads to speed up training
)

# Optional: Save the model for later use
# model.save("my_word2vec_model.model")

# Optional: Load a saved model
# model = Word2Vec.load("my_word2vec_model.model")

2. Printing Word-Level Similarity Scores

Now that we have a trained model, we can calculate similarity between individual words. For your two sentences, we can compare corresponding word pairs or find the most similar words across sentences.

Comparing Corresponding Word Pairs

Here's how to loop through your sentences and print similarity scores for each matching word pair:

sentence1 = "I am going to India"
sentence2 = "I am going to Bharat"

# Tokenize and lowercase to match model preprocessing
words1 = sentence1.lower().split()
words2 = sentence2.lower().split()

print("=== Word-Level Similarity Scores ===")
for w1, w2 in zip(words1, words2):
    try:
        # Fetch similarity score between the two words
        sim_score = model.wv.similarity(w1, w2)
        print(f"Similarity between '{w1}' and '{w2}': {sim_score:.4f}")
    except KeyError as e:
        print(f"Warning: Word '{e}' isn't in the model's vocabulary")

Finding Most Similar Words (for mismatched sentence lengths)

If your sentences don't have the same number of words, you can find the closest match in the second sentence for each word in the first:

print("\n=== Most Similar Words Across Sentences ===")
for w1 in words1:
    if w1 in model.wv:
        # Get the top 1 most similar word from the model's vocabulary
        most_similar = model.wv.most_similar(w1, topn=1)[0]
        print(f"'{w1}' is most similar to '{most_similar[0]}' with score: {most_similar[1]:.4f}")
    else:
        print(f"Word '{w1}' not found in vocabulary")

3. Calculating Sentence Similarity

Your original idea of using the average of word vectors is a solid baseline for sentence similarity. Let's fix your code to use the trained model's vectors correctly.

Step 1: Create Sentence Vectors (Average of Word Vectors)

def get_sentence_vector(sentence, model):
    words = sentence.lower().split()
    # Collect vectors only for words that exist in the model
    word_vectors = [model.wv[word] for word in words if word in model.wv]
    
    if not word_vectors:
        # Return a zero vector if no words are in the vocabulary
        return np.zeros(model.vector_size)
    
    # Calculate the average vector for the sentence
    return np.mean(word_vectors, axis=0)

# Generate vectors for your two sentences
sentence1_vec = get_sentence_vector(sentence1, model)
sentence2_vec = get_sentence_vector(sentence2, model)

Step 2: Calculate Cosine Similarity

You can implement this manually (like your original attempt) or use sklearn's built-in function for simplicity:

# Manual cosine similarity calculation
def cosine_similarity_manual(vec1, vec2):
    dot_product = np.dot(vec1, vec2)
    norm1 = np.linalg.norm(vec1)
    norm2 = np.linalg.norm(vec2)
    if norm1 == 0 or norm2 == 0:
        return 0.0
    return dot_product / (norm1 * norm2)

# Calculate and print results
manual_sim = cosine_similarity_manual(sentence1_vec, sentence2_vec)
sklearn_sim = cosine_similarity([sentence1_vec], [sentence2_vec])[0][0]

print("\n=== Sentence Similarity ===")
print(f"Manual calculation: {manual_sim:.4f}")
print(f"Sklearn calculation: {sklearn_sim:.4f}")

Full Working Code

Putting it all together, here's the complete script:

from gensim.models import Word2Vec
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity

# Train Word2Vec model
corpus = [
    "I am going to India",
    "I am going to Bharat",
    "India is a country in South Asia",
    "Bharat is the Hindi name for India",
    "People travel to India for holidays"
]
tokenized_corpus = [sent.lower().split() for sent in corpus]
model = Word2Vec(tokenized_corpus, vector_size=100, window=5, min_count=1, workers=4)

# Define target sentences
sentence1 = "I am going to India"
sentence2 = "I am going to Bharat"
words1 = sentence1.lower().split()
words2 = sentence2.lower().split()

# Print word-level similarities
print("=== Word-Level Similarity Scores ===")
for w1, w2 in zip(words1, words2):
    if w1 in model.wv and w2 in model.wv:
        sim = model.wv.similarity(w1, w2)
        print(f"'{w1}' vs '{w2}': {sim:.4f}")
    else:
        missing = [w for w in [w1, w2] if w not in model.wv]
        print(f"Missing words in vocabulary: {', '.join(missing)}")

# Calculate sentence similarity
def get_sentence_vec(sentence, model):
    vecs = [model.wv[w] for w in sentence.lower().split() if w in model.wv]
    return np.mean(vecs, axis=0) if vecs else np.zeros(model.vector_size)

vec1 = get_sentence_vec(sentence1, model)
vec2 = get_sentence_vec(sentence2, model)
sentence_sim = cosine_similarity([vec1], [vec2])[0][0]

print(f"\n=== Sentence Similarity ===")
print(f"Similarity between the two sentences: {sentence_sim:.4f}")

Quick Tips for You

  • Vocabulary Coverage: If a word isn't in your model's vocabulary, you can't get its vector. Training on more diverse text will improve this.
  • Preprocessing Consistency: Always align your input text with how you preprocessed the training corpus (lowercasing, tokenization, etc.).
  • Advanced Sentence Similarity: For better accuracy later, explore methods like Doc2Vec, BERT embeddings, or Sentence-BERT—but Word2Vec averages are a fantastic starting point.

内容的提问来源于stack exchange,提问作者marton mar suri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:58:19