You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用N-gram与Word2Vec时遭遇NameError:sentence_stream未定义

Fixing the NameError & Setting Up Bigram/Trigram Word2Vec

Hey there! Let's walk through solving that NameError first, then get your bigram/trigram-powered Word2Vec model up and running smoothly.

What's Causing the NameError?

The error NameError: name 'sentence_stream' is not defined happens because you referenced a variable sentence_stream that you never created. In your code, you already have a sentences variable (the one you passed to bigramer), so we'll use that as the base to generate our bigram and trigram streams.

Full Working Code

Here's a complete, corrected workflow that generates bigrams and trigrams, then trains Word2Vec with them. We'll stick to Gensim's Phrases and Phraser tools—they're optimized for integrating with Word2Vec, since they merge multi-word phrases into single tokens (like natural_language_processing) that the model can learn as cohesive units.

from gensim.models import Word2Vec, Phrases
from gensim.models.phrases import Phraser

# First, confirm your `sentences` is a nested list of tokenized words
# Replace this example with your actual dataset:
sentences = [["i", "love", "natural", "language", "processing"],
             ["word2vec", "is", "great", "for", "word", "embeddings"],
             ["bigrams", "and", "trigrams", "capture", "richer", "context"]]

# Step 1: Train a bigram model, then convert to lightweight Phraser for efficiency
bigram_model = Phrases(sentences, min_count=1, threshold=10)
bigram_phraser = Phraser(bigram_model)

# Step 2: Use bigram-processed sentences to train a trigram model
trigram_model = Phrases(bigram_phraser[sentences], min_count=1, threshold=10)
trigram_phraser = Phraser(trigram_model)

# Step 3: Train Word2Vec with trigram-enriched sentences
num_workers = 4
num_features = 100
min_word_count = 1
context_window = 5
downsampling_rate = 1e-3

word2vec_model = Word2Vec(
    trigram_phraser[bigram_phraser[sentences]],
    workers=num_workers,
    size=num_features,
    min_count=min_word_count,
    window=context_window,
    sample=downsampling_rate
)

# Optional: Lock the model to save memory after training
word2vec_model.init_sims(replace=True)

Key Tips for Your Project:

  • Phraser > raw Phrases: Phrases is the heavy training model, while Phraser is a lightweight, faster version meant for processing text. It's critical for efficiency with large datasets.
  • Tweak parameters for your data: Adjust min_count (ignore rare phrases) and threshold (higher values mean only very common, meaningful phrases are kept) based on how big your dataset is and what phrases you want to capture.
  • Chain processors: We first apply the bigram phraser to raw sentences, then feed that output into the trigram phraser. This ensures we build up multi-word phrases correctly.

内容的提问来源于stack exchange,提问作者beginner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:48:15