完成预处理后,如何用NLTK文本分析库预测特定文本(Python)
Alright, let's break this down step by step. You've already got the feature extraction part started with scikit-learn, but to integrate NLTK for prediction, we need to cover proper NLTK-based preprocessing first, then build a classifier and use it to predict on new text or text groups. Here's how to do it:
Your current code uses scikit-learn's tools, but NLTK provides more granular text cleaning steps that will improve model performance. Let's create a reusable preprocessing function:
import nltk from nltk.tokenize import word_tokenize from nltk.stem import WordNetLemmatizer from nltk.corpus import stopwords import string # Download required NLTK resources (run once) nltk.download('punkt') nltk.download('wordnet') nltk.download('stopwords') def preprocess_text(text): # Convert text to lowercase text = text.lower() # Split text into individual words (tokenization) tokens = word_tokenize(text) # Remove punctuation and non-alphabetic characters tokens = [token for token in tokens if token.isalpha()] # Filter out common stopwords (like "the", "and") stop_words = set(stopwords.words('english')) tokens = [token for token in tokens if token not in stop_words] # Reduce words to their base form (lemmatization) lemmatizer = WordNetLemmatizer() tokens = [lemmatizer.lemmatize(token) for token in tokens] # Join tokens back into a single string for feature extraction return ' '.join(tokens)
Apply this preprocessing to your entire corpus:
preprocessed_corpus = [preprocess_text(text) for text in corpus]
Replace your original corpus with the preprocessed version in your existing feature extraction pipeline:
from sklearn.feature_extraction.text import CountVectorizer, TfidfTransformer # Initialize vectorizer with your existing parameters vectorizer = CountVectorizer(max_features=2000, max_df=0.6, min_df=3) X_counts = vectorizer.fit_transform(preprocessed_corpus) # Convert word counts to TF-IDF scores transformer = TfidfTransformer() X_tfidf = transformer.fit_transform(X_counts)
NLTK's classifiers expect data in a specific (feature_dictionary, label) format. First, let's assume you have a label list y (length 2000, containing 'positive'/'negative' for each comment). Convert your TF-IDF features to NLTK's required format and train a classifier:
from nltk.classify import NaiveBayesClassifier def tfidf_to_nltk_format(X_tfidf, labels, vectorizer): feature_names = vectorizer.get_feature_names_out() nltk_training_data = [] for i in range(X_tfidf.shape[0]): # Get non-zero features for the current sample non_zero_indices = X_tfidf[i].nonzero()[1] feature_dict = {feature_names[idx]: X_tfidf[i, idx] for idx in non_zero_indices} nltk_training_data.append((feature_dict, labels[i])) return nltk_training_data # Prepare training data (replace `y` with your actual label list) nltk_train_data = tfidf_to_nltk_format(X_tfidf, y, vectorizer) # Train NLTK's Naive Bayes classifier classifier = NaiveBayesClassifier.train(nltk_train_data)
Now use the trained classifier to make predictions. Remember to apply the same preprocessing to new text:
def predict_sentiment(text, classifier, vectorizer, transformer): # Preprocess the input text preprocessed_text = preprocess_text(text) # Convert to word count vector using the trained vectorizer text_counts = vectorizer.transform([preprocessed_text]) # Convert to TF-IDF text_tfidf = transformer.transform(text_counts) # Convert to NLTK's feature dictionary format feature_names = vectorizer.get_feature_names_out() non_zero_indices = text_tfidf[0].nonzero()[1] features = {feature_names[idx]: text_tfidf[0, idx] for idx in non_zero_indices} # Return the predicted sentiment return classifier.classify(features) # Example: Predict a single comment new_comment = "This product exceeded my expectations, I'll definitely buy it again!" prediction = predict_sentiment(new_comment, classifier, vectorizer, transformer) print(f"Predicted sentiment: {prediction}") # Example: Predict a group of comments new_comments = [ "Worst purchase ever, it broke after one use.", "The service was great, and the product works as advertised.", "Meh, it's okay but not worth the price." ] for comment in new_comments: pred = predict_sentiment(comment, classifier, vectorizer, transformer) print(f"Comment: {comment}\nPredicted sentiment: {pred}\n")
If you prefer using scikit-learn's classifiers (like Logistic Regression) instead of NLTK's, you can integrate NLTK's preprocessing directly into the CountVectorizer:
def nltk_tokenizer(text): tokens = word_tokenize(text.lower()) tokens = [token for token in tokens if token.isalpha()] stop_words = set(stopwords.words('english')) tokens = [token for token in tokens if token not in stop_words] lemmatizer = WordNetLemmatizer() tokens = [lemmatizer.lemmatize(token) for token in tokens] return tokens # Initialize vectorizer with custom NLTK tokenizer vectorizer = CountVectorizer( tokenizer=nltk_tokenizer, max_features=2000, max_df=0.6, min_df=3, stop_words=None # We handle stopwords in the tokenizer ) X_counts = vectorizer.fit_transform(corpus) X_tfidf = transformer.fit_transform(X_counts) # Train a scikit-learn classifier from sklearn.linear_model import LogisticRegression clf = LogisticRegression() clf.fit(X_tfidf, y) # Predict function for scikit-learn def predict_with_sklearn(text, clf, vectorizer, transformer): text_counts = vectorizer.transform([text]) text_tfidf = transformer.transform(text_counts) return clf.predict(text_tfidf)[0] # Test it out print(predict_with_sklearn(new_comment, clf, vectorizer, transformer))
内容的提问来源于stack exchange,提问作者SHUBHENDRA KUMAR

