You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Word2Vec的句子相似度计算:代码运行失败求助

Hey there, fellow researcher! Let's troubleshoot that Word2Vec sentence similarity code you're working on. It looks like you've got a solid start, but there are a few common pitfalls that might be causing it to fail. Let's walk through fixing it step by step.

Fixing Your Word2Vec Sentence Similarity Code

First, let's address the incomplete part in your code snippet and fill in the gaps, plus fix the common issues that break execution:

1. Complete & Correct the Feature Vector Calculation

Your code cuts off mid-check for word membership in the vocabulary set. Here's the full, error-fixed version of the avg_feature_vector function:

import numpy as np
from scipy import spatial

def avg_feature_vector(sentence, model, num_features, index2word_set):
    words = sentence.split()
    feature_vec = np.zeros((num_features, ), dtype='float32')
    n_words = 0
    
    for word in words:
        # Fix: Complete the vocabulary check (you had a typo/cut-off here)
        if word in index2word_set:
            feature_vec += model.wv[word]
            n_words += 1
    
    # Handle edge case: no words from the sentence exist in the model's vocabulary
    if n_words > 0:
        feature_vec = feature_vec / n_words
    return feature_vec

2. Critical Pre-Requisites You Might Have Missed

Before running the function, make sure you've covered these often-overlooked steps:

  • Load/Train a Valid Word2Vec Model: You need either a custom-trained model on your research corpus or a pre-trained one (like Google's 300-dimensional Word2Vec model). Example code for loading:
    from gensim.models import Word2Vec
    from gensim.models.keyedvectors import KeyedVectors
    
    # For a custom-trained model:
    # model = Word2Vec.load("your_trained_model_path.model")
    # For a pre-trained binary model (e.g., Google's):
    # model = KeyedVectors.load_word2vec_format('GoogleNews-vectors-negative300.bin', binary=True)
    
  • Fix Vocabulary Set Initialization: Newer Gensim versions renamed index2word to index_to_key — use the correct attribute to avoid AttributeErrors:
    index2word_set = set(model.wv.index_to_key)  # Use model.wv.index2word for older Gensim versions
    
  • Match Feature Dimension: The num_features parameter must exactly match the dimensionality of your Word2Vec model (e.g., 300 for Google's pre-trained model).

3. Calculate Sentence Similarity

Once you have valid feature vectors for your sentences, use cosine similarity to compare them:

# Example usage with sample sentences
sentence1 = "Your first research sentence"
sentence2 = "Your second comparison sentence"

num_features = 300  # Match your model's dimension

# Generate average feature vectors for both sentences
vec1 = avg_feature_vector(sentence1, model, num_features, index2word_set)
vec2 = avg_feature_vector(sentence2, model, num_features, index2word_set)

# Compute cosine similarity (1 - distance = similarity score)
similarity_score = 1 - spatial.distance.cosine(vec1, vec2)
print(f"Sentence Similarity Score: {similarity_score:.4f}")

4. Common Error Fixes to Keep in Mind

  • KeyError for Unseen Words: The function skips words not in the model's vocabulary (which is intentional), but if a sentence has no overlapping words, it returns a zero vector. Add a warning for this edge case if needed.
  • Gensim Version Conflicts: If you get an AttributeError for index2word, switch to index_to_key (this is a breaking change in Gensim 4.0+).
  • Data Type Mismatches: Using float32 for the feature vector ensures compatibility with most Word2Vec model vector dtypes.

5. Bonus: Research-Grade Improvements

Since you're a grad student, here are tweaks to boost performance for your work:

  • Text Preprocessing: Add lowercasing, stopword removal, or lemmatization (using NLTK or spaCy) before splitting sentences to reduce noise.
  • Weighted Averages: Replace simple averaging with TF-IDF-weighted averages to prioritize semantically important words.
  • Alternative Approaches: If you can expand beyond basic Word2Vec, consider Doc2Vec (built for sentence-level embeddings) or Sentence-BERT for better accuracy — but the above method works perfectly if you need to stick to Word2Vec.

内容的提问来源于stack exchange,提问作者pintu lohar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:34:28