使用Gensim基于CSV训练Doc2vec模型时遇两类报错求助
Hey there! I’ve run into both these exact snags before, so let’s work through them one by one to get your model up and running.
1. Correctly Using TaggedDocument with Your Sr Field
The TaggedDocument class in Gensim needs two specific inputs: a list of tokenized words for the document, and a list of tags (not a single value or pandas Series directly). The generic 'tag' workaround works, but you can absolutely use your Sr field as unique identifiers—you just need to format it right.
Here’s the fix:
- Make sure each
Srvalue is wrapped in a list (even if it’s a single ID) - Convert
Srto a string if it’s numeric (avoids unexpected type issues)
Example code if you’re using a pandas DataFrame:
from gensim.models.doc2vec import TaggedDocument import pandas as pd # Sample data setup (match your actual DataFrame structure) df = pd.DataFrame({ 'comment_text': ['first unmeaningful comment', 'second random text', 'third sample comment'], 'Sr': [101, 102, 103] }) # Build your tagged corpus properly tagged_corpus = [ TaggedDocument(words=row['comment_text'].split(), tags=[str(row['Sr'])]) for _, row in df.iterrows() ]
Now each document gets its unique Sr value as a tag, which is way more useful than a generic label later on.
2. Fixing the ValueError During Training
That error happens because Gensim’s train() method needs to know the total size of your corpus to calculate progress and adjust training parameters correctly. The simplest solution is to pass total_examples=model.corpus_count—this value is automatically set when you build the vocabulary, so you don’t have to count documents manually.
Here’s how to update your training code:
from gensim.models import Doc2Vec # Initialize your model (adjust parameters to fit your data) model = Doc2Vec(vector_size=100, window=3, min_count=1, workers=4) # Build vocabulary from your tagged corpus model.build_vocab(tagged_corpus) # Train with the required parameter model.train( tagged_corpus, epochs=15, total_examples=model.corpus_count # This eliminates the ValueError! )
If you prefer, you could also calculate the total number of tokens in your corpus and pass total_words=your_total_token_count, but total_examples is almost always the easier, more reliable choice for Doc2Vec.
内容的提问来源于stack exchange,提问作者Shoaibkhanz

