Mallet LDA模型加载预测报[Errno 2]错误的解决咨询
Hey there! I get it, dealing with missing temporary files after saving a Gensim LdaMallet model is super frustrating—especially when you can load the model and see topics but can't run predictions. Let's break down the solutions here.
Why This Happens
When you trained your model without setting the prefix parameter, Gensim uses a random temporary directory (like the /var/folders/... path in your error) to store Mallet's intermediate files (like doctopics.txt.infer). These files aren't saved with the model, so when you load it later, Gensim can't find them to run inference.
Solution 1: Manually Calculate Document-Topic Probabilities (No Retraining Needed)
Since you can still access the topic-word weights from your loaded model, you can approximate document-topic probabilities using basic probability calculations. Here's how to do it:
Step 1: Extract Topic-Word Probability Distributions
First, we'll get the full probability distribution for each topic across your entire vocabulary:
import numpy as np # Load your existing model (use LdaMallet.load, not LdaModel.load!) model_lda = gensim.models.wrappers.LdaMallet.load('lda_v0.model') # Get topic-word probability arrays aligned with your id2word mapping topic_word_probs = [] vocab_size = len(model_lda.id2word) for topic_id in range(model_lda.num_topics): # Fetch all words and their probabilities for the topic topic_words = model_lda.show_topic(topic_id, topn=vocab_size) word_prob_map = {word: prob for word, prob in topic_words} # Create an array matching the order of id2word (add tiny epsilon to avoid log(0)) prob_array = np.array([word_prob_map.get(model_lda.id2word[word_id], 1e-10) for word_id in range(vocab_size)]) topic_word_probs.append(prob_array)
Step 2: Build a Prediction Function
Next, we'll create a function that takes a document (in corpus format: list of (word_id, count) tuples) and calculates its topic probabilities:
def predict_doc_topics(model, doc_corpus): # Convert document corpus to a frequency array doc_freqs = np.zeros(len(model.id2word)) for word_id, count in doc_corpus: doc_freqs[word_id] = count # Calculate unnormalized topic scores (incorporate alpha prior) alpha = model.alpha topic_scores = [] for topic_probs in topic_word_probs: # Sum log probabilities weighted by word frequency, plus prior log_score = np.sum(doc_freqs * np.log(topic_probs)) + np.log(alpha) topic_scores.append(log_score) # Normalize scores to probabilities using softmax (avoid overflow) topic_scores = np.array(topic_scores) exp_scores = np.exp(topic_scores - np.max(topic_scores)) doc_topic_probs = exp_scores / exp_scores.sum() # Return in Gensim's standard (topic_id, probability) format return [(topic_id, float(prob)) for topic_id, prob in enumerate(doc_topic_probs)]
How to Use It
Just pass your preprocessed input document (in the same corpus format you used for training) to the function:
# Example: input_corpus is your document in (word_id, count) format doc_topics = predict_doc_topics(model_lda, input_corpus) print(doc_topics)
⚠️ Note: This is an approximation—Mallet's native inference uses Gibbs sampling which is more accurate—but this method will give you reasonable topic probabilities if retraining isn't an option.
Solution 2: Retrain the Model with a Fixed prefix
If you still have access to your training data and can afford to retrain, this is the most reliable fix. By setting the prefix parameter, you tell Gensim to store Mallet's intermediate files in a permanent directory, which will be accessible when you load the model later.
mallet_path = 'mallet-2.0.8/bin/mallet' # Choose a permanent directory for Mallet's files (create it first if needed) prefix = './mallet_model_files/' # Train the model with the prefix set ldamallet = gensim.models.wrappers.LdaMallet( mallet_path, corpus=corpus, id2word=id2word, num_topics=14, prefix=prefix ) # Save the model and keep the prefix directory intact ldamallet.save('lda_v1.model')
Now when you load lda_v1.model later, Gensim will find all the necessary files in ./mallet_model_files/, and predictions will work as expected.
内容的提问来源于stack exchange,提问作者Omar Souaidi

