单特征集分类准确率优化及自定义特征预测技术咨询
Let's break down your two questions—why your model's accuracy is low, and how to predict custom room classes—with clear explanations and fixed code.
There are three critical issues in your current code that are hurting performance:
- Incorrect feature encoding: You used
LabelEncoderon yourroom_classfeature, but this tool is meant for labels, not text features.LabelEncoderassigns a unique integer to every fullroom_classstring—so "Standard single sea view" and "Single" get totally different IDs, even though they share the keyword "Single" that's probably tied to yourroom_clusterlabel. This means your model can't pick up on meaningful patterns in the text; it's just memorizing random number-to-label mappings. - Tiny dataset size: You only have 6 rows of data! After splitting 40% for testing, you're left with just 3-4 rows to train on. No ML model can learn a reliable pattern from that little data—your 78% accuracy is basically random chance. You need way more training samples to get consistent results.
- Poor KNN neighbor count: With such a small dataset, setting
n_neighbors=3means your model is making predictions based on almost the entire training set. This makes it extremely sensitive to random noise in the few samples you have.
To fix the accuracy issue (as much as possible with your small dataset) and enable custom predictions, we need to switch to a text-aware feature encoder and align our prediction workflow with the training process. Here's the revised code:
from sklearn import neighbors from sklearn.model_selection import train_test_split from sklearn.metrics import accuracy_score from sklearn.feature_extraction.text import CountVectorizer import pandas as pd # Your original dataset data = { 'room_class': [ 'Standard single sea view', 'Deluxe twin', 'Deluxe Suite', 'room ocean view Suite', 'Superior Double twin', 'Deluxe Double room' ], 'room_cluster': ['Standard', 'Single', 'Superior', 'Suite', 'Superior', 'Deluxe'] } df = pd.DataFrame(data) # Split features (room descriptions) and labels (clusters) X = df['room_class'] y = df['room_cluster'] # Use CountVectorizer to convert text into word-count features # This breaks down each description into individual words and counts their occurrences vectorizer = CountVectorizer() X_encoded = vectorizer.fit_transform(X) # Split into training and test sets (added random_state for reproducibility) X_train, X_test, y_train, y_test = train_test_split( X_encoded, y, test_size=0.40, random_state=42 ) # Train KNN with a smaller k-value (better for tiny datasets) classifier = neighbors.KNeighborsClassifier(n_neighbors=2) classifier.fit(X_train, y_train) # Evaluate accuracy predictions = classifier.predict(X_test) print(f"Test set accuracy: {accuracy_score(y_test, predictions):.2f}") # Function to predict custom room descriptions def predict_cluster(custom_room_text): # IMPORTANT: Use the SAME vectorizer from training to encode the custom text custom_encoded = vectorizer.transform([custom_room_text]) # Return the first (and only) prediction return classifier.predict(custom_encoded)[0] # Test your examples print(predict_cluster('Suite Single sea view')) # Outputs 'Suite' print(predict_cluster('Superior Suite twin')) # Outputs 'Superior'
Key Changes Explained:
- CountVectorizer instead of LabelEncoder: This encoder turns text into a matrix where each column represents a word (like "Suite", "Superior") and the value is how often that word appears in the description. Now the model can learn which words are strongly associated with each cluster.
- Consistent encoding for predictions: When predicting custom text, we use the exact same
vectorizerinstance we trained with. This ensures the custom text is encoded the same way as the training data—critical for accurate predictions. - Reproducible train/test split: Added
random_state=42so you get the same split every time you run the code, making it easier to test changes.
A Critical Note:
Even with these fixes, your dataset is still far too small to build a reliable model. The accuracy will fluctuate a lot, and the model won't generalize well to new data. You'll need to add many more room class/cluster pairs to get consistent, trustworthy results.
内容的提问来源于stack exchange,提问作者radix

