You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gensim Word2Vec most similar结果异常求助:基于哈利波特文本的Python实验

Hey Camilla, let's figure out why you're getting inconsistent results from most_similar after processing your Harry Potter text, and work through solutions to fix this:

Possible Causes & Fixes

1. Control Randomness in Word2Vec Training

Word2Vec has several random elements under the hood—like weight initialization, negative sampling selection, and training order shuffling. Even with the same corpus, these can lead to different model outputs across runs. If your "different results" come from repeated training or mismatches between Hermione_1 and Hermione_2 vectors, fix this by locking in random seeds:

import random
import numpy as np
from gensim.models import Word2Vec

# Fix global random seeds
random.seed(42)
np.random.seed(42)

# Pass seed to Word2Vec during training
model = Word2Vec(
    your_corpus,
    vector_size=100,
    window=5,
    min_count=1,
    workers=4,
    seed=42  # Critical for reproducibility
)

With fixed seeds, you’ll get identical vector outputs every time you train on the same corpus, making it easy to compare Hermione_1 and Hermione_2 consistently.

2. Ensure Hermione_1 and Hermione_2 Have Identical Contexts

Your workflow replaces Hermione with two variants in separate files, then combines them. If the replacements aren’t 100% consistent (e.g., missing instances with punctuation like Hermione. or Hermione,), the two tokens will have different context distributions—leading to different similarity results.

First, Verify Replacement Consistency

Use this quick script to check if Hermione_1 and Hermione_2 share the exact same contexts:

from collections import defaultdict

def collect_contexts(file_path, target_token, window=5):
    context_counts = defaultdict(int)
    with open(file_path, 'r') as f:
        for line in f:
            tokens = line.strip().split()
            for idx, token in enumerate(tokens):
                if token == target_token:
                    # Capture window of words around the target
                    left_context = tokens[max(0, idx - window):idx]
                    right_context = tokens[idx + 1:idx + 1 + window]
                    context_key = tuple(left_context + right_context)
                    context_counts[context_key] += 1
    return context_counts

# Compare contexts for both tokens
ctx1 = collect_contexts("Hermione_1.txt", "Hermione_1")
ctx2 = collect_contexts("Hermione_2.txt", "Hermione_2")

print("Contexts match:", ctx1 == ctx2)

If this prints False, your replacement logic is incomplete. Fix it to handle all occurrences of Hermione, including those attached to punctuation:

import re

def safe_replace_hermione(input_file, output_file, replacement):
    with open(input_file, 'r') as in_f, open(output_file, 'w') as out_f:
        for line in in_f:
            # Match standalone Hermione (ignoring surrounding punctuation)
            updated_line = re.sub(r'\bHermione\b', replacement, line)
            out_f.write(updated_line)

# Regenerate your two files with this improved function
safe_replace_hermione("HarryPotter1.txt", "Hermione_1.txt", "Hermione_1")
safe_replace_hermione("HarryPotter1.txt", "Hermione_2.txt", "Hermione_2")

Now Hermione_1 and Hermione_2 will have identical context patterns, so their vectors should be nearly identical (exact matches with fixed seeds).

3. Tune Word2Vec Parameters for Stability

A single Harry Potter book might not be a huge corpus, which can make Word2Vec results unstable. Adjust these parameters to improve vector reliability:

  • vector_size: Increase from 100 to 200/300 to give the model more space to capture nuance
  • window: Widen from 5 to 10 to consider broader context around each token
  • epochs: Boost from the default 5 to 10-15 to let the model learn the corpus more thoroughly
  • min_count: Keep it low (1) if Hermione appears frequently enough

Example adjusted training code:

model = Word2Vec(
    your_corpus,
    vector_size=200,
    window=10,
    min_count=1,
    workers=4,
    seed=42,
    epochs=15
)

4. Double-Check Corpus Concatenation

Make sure you’re combining the two files correctly without introducing encoding issues or missing content. A safe way to concatenate is:

with open("combined_corpus.txt", 'w', encoding='utf-8') as combined_f:
    for file_name in ["Hermione_1.txt", "Hermione_2.txt"]:
        with open(file_name, 'r', encoding='utf-8') as f:
            combined_f.write(f.read())

Using explicit UTF-8 encoding avoids garbled text that could throw off training.

Verify Your Fix

Once you’ve implemented these changes, confirm the vectors are consistent with:

similarity_score = model.wv.similarity("Hermione_1", "Hermione_2")
print(f"Hermione_1 vs Hermione_2 similarity: {similarity_score:.4f}")

With proper fixes, this score should be very close to 1.0 (exactly 1.0 if all variables are controlled perfectly).

内容的提问来源于stack exchange,提问作者Camilla8

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:36:33