基于监督机器学习的Python推特文本分类实现技术问询
Great question! Given your small labeled dataset (20 samples per class), the key here is to prioritize robust text preprocessing and choose models that perform well with limited data. Below is a step-by-step Python implementation using scikit-learn, which is straightforward and effective for this scenario:
1. Setup Dependencies
First, install the required libraries if you haven't already:
pip install pandas scikit-learn nltk
Then import them in your script:
import pandas as pd import re from nltk.corpus import stopwords from nltk.stem import WordNetLemmatizer from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.model_selection import train_test_split from sklearn.naive_bayes import MultinomialNB from sklearn.svm import LinearSVC from sklearn.metrics import accuracy_score, classification_report # Download NLTK resources (run once) import nltk nltk.download('stopwords') nltk.download('wordnet')
2. Load and Preprocess Data
First, load your labeled and unlabeled datasets (assuming they're in CSV format with columns text and label for labeled data, just text for unlabeled):
# Load labeled data labeled_df = pd.read_csv('labeled_tweets.csv') # Load unlabeled data unlabeled_df = pd.read_csv('unlabeled_tweets.csv')
Next, create a text preprocessing function tailored to Twitter content:
def preprocess_text(text): # Convert to lowercase text = text.lower() # Remove URLs text = re.sub(r'https?://\S+|www\.\S+', '', text) # Remove mentions (@username) text = re.sub(r'@\w+', '', text) # Remove special characters and numbers text = re.sub(r'[^a-zA-Z\s]', '', text) # Tokenize and remove stopwords stop_words = set(stopwords.words('english')) words = text.split() words = [word for word in words if word not in stop_words] # Lemmatization to reduce words to their base form lemmatizer = WordNetLemmatizer() words = [lemmatizer.lemmatize(word) for word in words] # Join back to a single string return ' '.join(words) # Apply preprocessing to both datasets labeled_df['clean_text'] = labeled_df['text'].apply(preprocess_text) unlabeled_df['clean_text'] = unlabeled_df['text'].apply(preprocess_text)
3. Feature Extraction with TF-IDF
Convert text to numerical features using TF-IDF, which is a standard and effective method for text classification:
# Initialize TF-IDF Vectorizer (limit features to avoid overfitting) tfidf = TfidfVectorizer(max_features=1000) # Fit on labeled data and transform both labeled and unlabeled text X_labeled = tfidf.fit_transform(labeled_df['clean_text']) y_labeled = labeled_df['label'] X_unlabeled = tfidf.transform(unlabeled_df['clean_text'])
4. Train and Evaluate Models
With a small dataset, Multinomial Naive Bayes and Linear SVM are top choices—they’re lightweight, fast, and perform well on text data. Let’s train both and compare their performance:
Option 1: Multinomial Naive Bayes
# Split labeled data into train/test sets (80-20 split) X_train, X_test, y_train, y_test = train_test_split(X_labeled, y_labeled, test_size=0.2, random_state=42) # Train the model nb_model = MultinomialNB() nb_model.fit(X_train, y_train) # Evaluate performance y_pred = nb_model.predict(X_test) print("Naive Bayes Accuracy:", accuracy_score(y_test, y_pred)) print("\nClassification Report:\n", classification_report(y_test, y_pred))
Option 2: Linear SVM
# Train the model svm_model = LinearSVC() svm_model.fit(X_train, y_train) # Evaluate performance y_pred_svm = svm_model.predict(X_test) print("Linear SVM Accuracy:", accuracy_score(y_test, y_pred_svm)) print("\nClassification Report:\n", classification_report(y_test, y_pred_svm))
Linear SVM often outperforms Naive Bayes on text tasks, but test both to see which works better for your specific data.
5. Predict on Unlabeled Data
Once you’ve selected your best model, apply it to classify the remaining unlabeled tweets:
# Use the SVM model as an example (replace with nb_model if preferred) unlabeled_df['predicted_label'] = svm_model.predict(X_unlabeled) # Save the classified results to a CSV file unlabeled_df.to_csv('classified_tweets.csv', index=False)
Optional Improvements for Small Datasets
Since you only have 20 samples per class, here are some ways to boost model performance:
- Data Augmentation: Use tools like
nlpaugto generate additional labeled samples via back-translation or synonym replacement. - Hyperparameter Tuning: Use
GridSearchCVto optimize model parameters (e.g.,alphafor Naive Bayes,Cfor SVM). - Pre-trained Language Models: For better accuracy (though more compute-heavy), use models like BERT via Hugging Face’s
transformerslibrary—pre-trained models leverage existing language knowledge to perform well with limited labeled data.
内容的提问来源于stack exchange,提问作者Natalie

