基于Scikit-learn SVM训练分类器:未见过数据分类求助
Alright, here's a clean, structured breakdown of your text classification workflow formatted properly with Markdown:
Text Classification Workflow Summary
1. Sample Dataset
Below is your labeled training data presented in a readable table:
| ID | Comment | Category |
|---|---|---|
| 2017_01 | inadequate stock | Availability |
| 2017_02 | Too many failures | Quality |
| 2017_03 | no documentation | Customer Service |
| 2017_04 | good product | Satisfied |
| 2017_05 | long delivery times | Delivery |
2. Model Selection & Implementation
You tested two standard multi-class text classifiers and settled on Support Vector Machines (SVM) over MultinomialNB for better fit. Here’s a practical code snippet to implement and fit the SVM model using scikit-learn:
from sklearn.svm import SVC from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.pipeline import Pipeline from sklearn.model_selection import train_test_split # Prepare your dataset (adjust based on how you're storing your data) X = ["inadequate stock", "Too many failures", "no documentation", "good product", "long delivery times"] y = ["Availability", "Quality", "Customer Service", "Satisfied", "Delivery"] # Split into training/validation sets (good practice for evaluation) X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42) # Build a end-to-end pipeline: TF-IDF text vectorization + SVM classifier text_classifier = Pipeline([ ('tfidf', TfidfVectorizer(stop_words='english')), # Remove common stopwords for cleaner features ('svm', SVC(kernel='linear', C=1.0)) # Linear kernel is generally effective for text tasks ]) # Train the model on your labeled data text_classifier.fit(X_train, y_train) # Classify unseen data with the trained model unseen_comments = ["my order took 2 weeks to arrive", "item stopped working after 3 days"] predicted_categories = text_classifier.predict(unseen_comments) print(predicted_categories)
Pro tip: To optimize performance, consider tuning hyperparameters like
C(regularization strength for SVM) or n-gram ranges in TF-IDF using cross-validation.
内容的提问来源于stack exchange,提问作者Chat Peters
相关产品推荐
相关产品推荐

