如何将训练好的NLTK NaiveBayesClassifier应用于文本分类?
Great job getting your NaiveBayesClassifier trained! Now let's walk through exactly how to apply it to your text file of sentences. The key here is to reuse the exact same preprocessing pipeline you used during training—this ensures the classifier gets features it recognizes.
Step 1: Load Your Trained Classifier
First, you need to load the classifier you saved (I assume you used pickle for this, which is standard for NLTK models). If you haven't saved it yet, do that first with:
import pickle # Save the classifier (run this once after training) with open('question_classifier.pkl', 'wb') as f: pickle.dump(classifier, f)
Then load it when you're ready to classify:
import pickle from nltk.classify import NaiveBayesClassifier # Load the trained classifier with open('question_classifier.pkl', 'rb') as f: classifier = pickle.load(f)
Step 2: Define Your Preprocessing Function
This is critical—your test sentences need the same treatment as your training data. Let's assume your training preprocessing included tokenization, stemming, stopword removal, and creating a bag-of-words feature set. Here's a reusable function for that:
from nltk.tokenize import word_tokenize from nltk.stem import PorterStemmer from nltk.corpus import stopwords import string stop_words = set(stopwords.words('english') + list(string.punctuation)) stemmer = PorterStemmer() def preprocess_sentence(sentence): # Tokenize the sentence tokens = word_tokenize(sentence.lower()) # Remove stopwords and punctuation, then stem filtered_tokens = [stemmer.stem(token) for token in tokens if token not in stop_words] # Create bag-of-words features return dict([(token, True) for token in filtered_tokens])
Adjust this function to match exactly what you did during training (e.g., if you used a different stemmer or stopword list, swap those out). For banking-specific text, you might even want to add domain-specific stopwords like "account" or "bank" if they don't add classification value.
Step 3: Process Your Text File
Now read each line from your text file, preprocess it, and classify:
# Path to your input text file input_file_path = 'bank_sentences.txt' # Optional: Path to save results output_file_path = 'classified_bank_sentences.txt' with open(input_file_path, 'r', encoding='utf-8') as infile, open(output_file_path, 'w', encoding='utf-8') as outfile: for line in infile: sentence = line.strip() if not sentence: # Skip empty lines continue # Preprocess to get features features = preprocess_sentence(sentence) # Add a quick question mark check (optional but helpful) features['has_question_mark'] = '?' in sentence # Classify the sentence classification = classifier.classify(features) # Write the result (customize the format as needed) outfile.write(f"Original: {sentence}\nClass: {classification}\n\n") print("Classification complete! Results saved to", output_file_path)
Step 4: Verify and Adjust
After running the code, spot-check a few sentences to make sure the classifications make sense. If you notice errors, you might need to:
- Tweak your preprocessing (e.g., refine stopwords or add more features like sentence length)
- Add more training data specific to banking text to improve accuracy
- Validate if your feature set is capturing meaningful signals (the question mark check is a simple but effective addition for obvious questions)
内容的提问来源于stack exchange,提问作者PolkaDot

