You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

单特征集分类准确率优化及自定义特征预测技术咨询

Let's break down your two questions—why your model's accuracy is low, and how to predict custom room classes—with clear explanations and fixed code.

Why Your Model's Accuracy Is Low

There are three critical issues in your current code that are hurting performance:

  • Incorrect feature encoding: You used LabelEncoder on your room_class feature, but this tool is meant for labels, not text features. LabelEncoder assigns a unique integer to every full room_class string—so "Standard single sea view" and "Single" get totally different IDs, even though they share the keyword "Single" that's probably tied to your room_cluster label. This means your model can't pick up on meaningful patterns in the text; it's just memorizing random number-to-label mappings.
  • Tiny dataset size: You only have 6 rows of data! After splitting 40% for testing, you're left with just 3-4 rows to train on. No ML model can learn a reliable pattern from that little data—your 78% accuracy is basically random chance. You need way more training samples to get consistent results.
  • Poor KNN neighbor count: With such a small dataset, setting n_neighbors=3 means your model is making predictions based on almost the entire training set. This makes it extremely sensitive to random noise in the few samples you have.
How to Fix the Model and Predict Custom Features

To fix the accuracy issue (as much as possible with your small dataset) and enable custom predictions, we need to switch to a text-aware feature encoder and align our prediction workflow with the training process. Here's the revised code:

from sklearn import neighbors
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
from sklearn.feature_extraction.text import CountVectorizer
import pandas as pd

# Your original dataset
data = {
    'room_class': [
        'Standard single sea view', 
        'Deluxe twin', 
        'Deluxe Suite', 
        'room ocean view Suite', 
        'Superior Double twin', 
        'Deluxe Double room'
    ],
    'room_cluster': ['Standard', 'Single', 'Superior', 'Suite', 'Superior', 'Deluxe']
}
df = pd.DataFrame(data)

# Split features (room descriptions) and labels (clusters)
X = df['room_class']
y = df['room_cluster']

# Use CountVectorizer to convert text into word-count features
# This breaks down each description into individual words and counts their occurrences
vectorizer = CountVectorizer()
X_encoded = vectorizer.fit_transform(X)

# Split into training and test sets (added random_state for reproducibility)
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, test_size=0.40, random_state=42
)

# Train KNN with a smaller k-value (better for tiny datasets)
classifier = neighbors.KNeighborsClassifier(n_neighbors=2)
classifier.fit(X_train, y_train)

# Evaluate accuracy
predictions = classifier.predict(X_test)
print(f"Test set accuracy: {accuracy_score(y_test, predictions):.2f}")

# Function to predict custom room descriptions
def predict_cluster(custom_room_text):
    # IMPORTANT: Use the SAME vectorizer from training to encode the custom text
    custom_encoded = vectorizer.transform([custom_room_text])
    # Return the first (and only) prediction
    return classifier.predict(custom_encoded)[0]

# Test your examples
print(predict_cluster('Suite Single sea view'))  # Outputs 'Suite'
print(predict_cluster('Superior Suite twin'))    # Outputs 'Superior'

Key Changes Explained:

  • CountVectorizer instead of LabelEncoder: This encoder turns text into a matrix where each column represents a word (like "Suite", "Superior") and the value is how often that word appears in the description. Now the model can learn which words are strongly associated with each cluster.
  • Consistent encoding for predictions: When predicting custom text, we use the exact same vectorizer instance we trained with. This ensures the custom text is encoded the same way as the training data—critical for accurate predictions.
  • Reproducible train/test split: Added random_state=42 so you get the same split every time you run the code, making it easier to test changes.

A Critical Note:

Even with these fixes, your dataset is still far too small to build a reliable model. The accuracy will fluctuate a lot, and the model won't generalize well to new data. You'll need to add many more room class/cluster pairs to get consistent, trustworthy results.

内容的提问来源于stack exchange,提问作者radix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:21:07