You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于监督机器学习的Python推特文本分类实现技术问询

Supervised ML Implementation for Twitter Text Classification (Small Labeled Dataset)

Great question! Given your small labeled dataset (20 samples per class), the key here is to prioritize robust text preprocessing and choose models that perform well with limited data. Below is a step-by-step Python implementation using scikit-learn, which is straightforward and effective for this scenario:

1. Setup Dependencies

First, install the required libraries if you haven't already:

pip install pandas scikit-learn nltk

Then import them in your script:

import pandas as pd
import re
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.svm import LinearSVC
from sklearn.metrics import accuracy_score, classification_report

# Download NLTK resources (run once)
import nltk
nltk.download('stopwords')
nltk.download('wordnet')

2. Load and Preprocess Data

First, load your labeled and unlabeled datasets (assuming they're in CSV format with columns text and label for labeled data, just text for unlabeled):

# Load labeled data
labeled_df = pd.read_csv('labeled_tweets.csv')
# Load unlabeled data
unlabeled_df = pd.read_csv('unlabeled_tweets.csv')

Next, create a text preprocessing function tailored to Twitter content:

def preprocess_text(text):
    # Convert to lowercase
    text = text.lower()
    # Remove URLs
    text = re.sub(r'https?://\S+|www\.\S+', '', text)
    # Remove mentions (@username)
    text = re.sub(r'@\w+', '', text)
    # Remove special characters and numbers
    text = re.sub(r'[^a-zA-Z\s]', '', text)
    # Tokenize and remove stopwords
    stop_words = set(stopwords.words('english'))
    words = text.split()
    words = [word for word in words if word not in stop_words]
    # Lemmatization to reduce words to their base form
    lemmatizer = WordNetLemmatizer()
    words = [lemmatizer.lemmatize(word) for word in words]
    # Join back to a single string
    return ' '.join(words)

# Apply preprocessing to both datasets
labeled_df['clean_text'] = labeled_df['text'].apply(preprocess_text)
unlabeled_df['clean_text'] = unlabeled_df['text'].apply(preprocess_text)

3. Feature Extraction with TF-IDF

Convert text to numerical features using TF-IDF, which is a standard and effective method for text classification:

# Initialize TF-IDF Vectorizer (limit features to avoid overfitting)
tfidf = TfidfVectorizer(max_features=1000)

# Fit on labeled data and transform both labeled and unlabeled text
X_labeled = tfidf.fit_transform(labeled_df['clean_text'])
y_labeled = labeled_df['label']
X_unlabeled = tfidf.transform(unlabeled_df['clean_text'])

4. Train and Evaluate Models

With a small dataset, Multinomial Naive Bayes and Linear SVM are top choices—they’re lightweight, fast, and perform well on text data. Let’s train both and compare their performance:

Option 1: Multinomial Naive Bayes

# Split labeled data into train/test sets (80-20 split)
X_train, X_test, y_train, y_test = train_test_split(X_labeled, y_labeled, test_size=0.2, random_state=42)

# Train the model
nb_model = MultinomialNB()
nb_model.fit(X_train, y_train)

# Evaluate performance
y_pred = nb_model.predict(X_test)
print("Naive Bayes Accuracy:", accuracy_score(y_test, y_pred))
print("\nClassification Report:\n", classification_report(y_test, y_pred))

Option 2: Linear SVM

# Train the model
svm_model = LinearSVC()
svm_model.fit(X_train, y_train)

# Evaluate performance
y_pred_svm = svm_model.predict(X_test)
print("Linear SVM Accuracy:", accuracy_score(y_test, y_pred_svm))
print("\nClassification Report:\n", classification_report(y_test, y_pred_svm))

Linear SVM often outperforms Naive Bayes on text tasks, but test both to see which works better for your specific data.

5. Predict on Unlabeled Data

Once you’ve selected your best model, apply it to classify the remaining unlabeled tweets:

# Use the SVM model as an example (replace with nb_model if preferred)
unlabeled_df['predicted_label'] = svm_model.predict(X_unlabeled)

# Save the classified results to a CSV file
unlabeled_df.to_csv('classified_tweets.csv', index=False)

Optional Improvements for Small Datasets

Since you only have 20 samples per class, here are some ways to boost model performance:

  • Data Augmentation: Use tools like nlpaug to generate additional labeled samples via back-translation or synonym replacement.
  • Hyperparameter Tuning: Use GridSearchCV to optimize model parameters (e.g., alpha for Naive Bayes, C for SVM).
  • Pre-trained Language Models: For better accuracy (though more compute-heavy), use models like BERT via Hugging Face’s transformers library—pre-trained models leverage existing language knowledge to perform well with limited labeled data.

内容的提问来源于stack exchange,提问作者Natalie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:02:47