Gensim Word2Vec most similar结果异常求助:基于哈利波特文本的Python实验
Hey Camilla, let's figure out why you're getting inconsistent results from most_similar after processing your Harry Potter text, and work through solutions to fix this:
1. Control Randomness in Word2Vec Training
Word2Vec has several random elements under the hood—like weight initialization, negative sampling selection, and training order shuffling. Even with the same corpus, these can lead to different model outputs across runs. If your "different results" come from repeated training or mismatches between Hermione_1 and Hermione_2 vectors, fix this by locking in random seeds:
import random import numpy as np from gensim.models import Word2Vec # Fix global random seeds random.seed(42) np.random.seed(42) # Pass seed to Word2Vec during training model = Word2Vec( your_corpus, vector_size=100, window=5, min_count=1, workers=4, seed=42 # Critical for reproducibility )
With fixed seeds, you’ll get identical vector outputs every time you train on the same corpus, making it easy to compare Hermione_1 and Hermione_2 consistently.
2. Ensure Hermione_1 and Hermione_2 Have Identical Contexts
Your workflow replaces Hermione with two variants in separate files, then combines them. If the replacements aren’t 100% consistent (e.g., missing instances with punctuation like Hermione. or Hermione,), the two tokens will have different context distributions—leading to different similarity results.
First, Verify Replacement Consistency
Use this quick script to check if Hermione_1 and Hermione_2 share the exact same contexts:
from collections import defaultdict def collect_contexts(file_path, target_token, window=5): context_counts = defaultdict(int) with open(file_path, 'r') as f: for line in f: tokens = line.strip().split() for idx, token in enumerate(tokens): if token == target_token: # Capture window of words around the target left_context = tokens[max(0, idx - window):idx] right_context = tokens[idx + 1:idx + 1 + window] context_key = tuple(left_context + right_context) context_counts[context_key] += 1 return context_counts # Compare contexts for both tokens ctx1 = collect_contexts("Hermione_1.txt", "Hermione_1") ctx2 = collect_contexts("Hermione_2.txt", "Hermione_2") print("Contexts match:", ctx1 == ctx2)
If this prints False, your replacement logic is incomplete. Fix it to handle all occurrences of Hermione, including those attached to punctuation:
import re def safe_replace_hermione(input_file, output_file, replacement): with open(input_file, 'r') as in_f, open(output_file, 'w') as out_f: for line in in_f: # Match standalone Hermione (ignoring surrounding punctuation) updated_line = re.sub(r'\bHermione\b', replacement, line) out_f.write(updated_line) # Regenerate your two files with this improved function safe_replace_hermione("HarryPotter1.txt", "Hermione_1.txt", "Hermione_1") safe_replace_hermione("HarryPotter1.txt", "Hermione_2.txt", "Hermione_2")
Now Hermione_1 and Hermione_2 will have identical context patterns, so their vectors should be nearly identical (exact matches with fixed seeds).
3. Tune Word2Vec Parameters for Stability
A single Harry Potter book might not be a huge corpus, which can make Word2Vec results unstable. Adjust these parameters to improve vector reliability:
vector_size: Increase from 100 to 200/300 to give the model more space to capture nuancewindow: Widen from 5 to 10 to consider broader context around each tokenepochs: Boost from the default 5 to 10-15 to let the model learn the corpus more thoroughlymin_count: Keep it low (1) ifHermioneappears frequently enough
Example adjusted training code:
model = Word2Vec( your_corpus, vector_size=200, window=10, min_count=1, workers=4, seed=42, epochs=15 )
4. Double-Check Corpus Concatenation
Make sure you’re combining the two files correctly without introducing encoding issues or missing content. A safe way to concatenate is:
with open("combined_corpus.txt", 'w', encoding='utf-8') as combined_f: for file_name in ["Hermione_1.txt", "Hermione_2.txt"]: with open(file_name, 'r', encoding='utf-8') as f: combined_f.write(f.read())
Using explicit UTF-8 encoding avoids garbled text that could throw off training.
Verify Your Fix
Once you’ve implemented these changes, confirm the vectors are consistent with:
similarity_score = model.wv.similarity("Hermione_1", "Hermione_2") print(f"Hermione_1 vs Hermione_2 similarity: {similarity_score:.4f}")
With proper fixes, this score should be very close to 1.0 (exactly 1.0 if all variables are controlled perfectly).
内容的提问来源于stack exchange,提问作者Camilla8

