Word2Vec/Doc2Vec训练失败:输入样本数与预期不符求助
Hey there, let's dig into this warning you're seeing. That "supplied example count (0) did not equal expected count (2381)" message tells us that when Gensim tried to train the model, it couldn't read any valid sentences from your corpus—even though you confirmed the corpus has 2381 samples. Here are the most likely fixes to get your training working:
1. Double-check your corpus format
Gensim's Word2Vec expects an iterable of lists of strings—each element in your corpus should be a list of words (e.g., [["hello", "world"], ["foo", "bar"]]), not raw strings or other formats. If your corpus is a list of full sentences (like ["hello world", "foo bar"]), you need to split each sentence into individual words first:
# Convert raw string sentences to word lists corpus = [sentence.split() for sentence in your_original_corpus]
Also, make sure there are no empty sentences in your corpus—Gensim skips these, which could skew the count if you have a lot of them. You can validate this with:
for idx, sent in enumerate(corpus): if not isinstance(sent, list) or len(sent) == 0: print(f"Problem at index {idx}: {sent}")
2. Ensure your corpus is reusable
If you used a generator to create your corpus (instead of a list), you might have already exhausted it when you counted the 2381 samples. Generators only iterate once—so when you pass it to train(), it returns nothing. Fix this by converting your corpus to a list first:
# Convert generator to a reusable list corpus = list(your_corpus_generator) # Now verify the length is still 2381 assert len(corpus) == 2381, "Corpus length mismatch after conversion!"
3. Verify corpus_count matches your expected number
After running build_vocab(), print out word2vec_model.corpus_count—it should be 2381. If it's not, that means build_vocab() didn't process your corpus correctly (likely due to format issues). This value is what train() uses as the expected count, so a mismatch here guarantees the warning.
4. Simplify to isolate the issue
Temporarily remove the intersect_word2vec_format() step and run training. If the warning goes away, the problem might be related to how the pre-trained vectors are interacting with your vocab (though this is less likely). If it still appears, the issue is definitely with your corpus or how you're passing it to train().
5. Adjust for Gensim version differences
- For Gensim 3.6: Stick with passing
total_examples=word2vec_model.corpus_count, but make sure the corpus passed totrain()is the exact same one used inbuild_vocab(). - For Gensim 4.3: The
train()method has been streamlined—you can omittotal_examplesentirely, since the model tracks this automatically. Try:word2vec_model.train(corpus, epochs=15)
Fixed code example (Gensim 3.6)
Here's a cleaned-up version of your code with validation steps:
# First, ensure corpus is in the correct format (list of word lists) corpus = [sent.split() for sent in your_raw_corpus] # Adjust this to your actual data source assert len(corpus) == 2381, "Corpus doesn't have expected 2381 samples!" word2vec_model = Word2Vec(size=300, window=5, min_count=2, workers=-1) word2vec_model.build_vocab(corpus) print(f"Corpus count after build_vocab: {word2vec_model.corpus_count}") # Should output 2381 # Load pre-trained vectors (optional, but keep it if you need it) word2vec_model.intersect_word2vec_format('GoogleNews-vectors-negative300.bin.gz', lockf=1.0, binary=True) # Train with the same valid corpus word2vec_model.train(corpus, total_examples=word2vec_model.corpus_count, epochs=15)
The root cause here is almost always a corpus that's either in the wrong format, has been exhausted (if using a generator), or contains invalid/empty sentences. Once you fix that, the warning should disappear and your model will train properly.
内容的提问来源于stack exchange,提问作者VJ.BG

