使用N-gram与Word2Vec时遭遇NameError:sentence_stream未定义
Hey there! Let's walk through solving that NameError first, then get your bigram/trigram-powered Word2Vec model up and running smoothly.
What's Causing the NameError?
The error NameError: name 'sentence_stream' is not defined happens because you referenced a variable sentence_stream that you never created. In your code, you already have a sentences variable (the one you passed to bigramer), so we'll use that as the base to generate our bigram and trigram streams.
Full Working Code
Here's a complete, corrected workflow that generates bigrams and trigrams, then trains Word2Vec with them. We'll stick to Gensim's Phrases and Phraser tools—they're optimized for integrating with Word2Vec, since they merge multi-word phrases into single tokens (like natural_language_processing) that the model can learn as cohesive units.
from gensim.models import Word2Vec, Phrases from gensim.models.phrases import Phraser # First, confirm your `sentences` is a nested list of tokenized words # Replace this example with your actual dataset: sentences = [["i", "love", "natural", "language", "processing"], ["word2vec", "is", "great", "for", "word", "embeddings"], ["bigrams", "and", "trigrams", "capture", "richer", "context"]] # Step 1: Train a bigram model, then convert to lightweight Phraser for efficiency bigram_model = Phrases(sentences, min_count=1, threshold=10) bigram_phraser = Phraser(bigram_model) # Step 2: Use bigram-processed sentences to train a trigram model trigram_model = Phrases(bigram_phraser[sentences], min_count=1, threshold=10) trigram_phraser = Phraser(trigram_model) # Step 3: Train Word2Vec with trigram-enriched sentences num_workers = 4 num_features = 100 min_word_count = 1 context_window = 5 downsampling_rate = 1e-3 word2vec_model = Word2Vec( trigram_phraser[bigram_phraser[sentences]], workers=num_workers, size=num_features, min_count=min_word_count, window=context_window, sample=downsampling_rate ) # Optional: Lock the model to save memory after training word2vec_model.init_sims(replace=True)
Key Tips for Your Project:
- Phraser > raw Phrases:
Phrasesis the heavy training model, whilePhraseris a lightweight, faster version meant for processing text. It's critical for efficiency with large datasets. - Tweak parameters for your data: Adjust
min_count(ignore rare phrases) andthreshold(higher values mean only very common, meaningful phrases are kept) based on how big your dataset is and what phrases you want to capture. - Chain processors: We first apply the bigram phraser to raw sentences, then feed that output into the trigram phraser. This ensures we build up multi-word phrases correctly.
内容的提问来源于stack exchange,提问作者beginner

