每次运行K近邻分类器代码准确率不同,求原因解析
Hey Ashish, I totally get why this is confusing—let's walk through exactly what's happening here.
The Root Cause: Random Train-Test Splitting
The main issue is how train_test_split works by default. Every time you run your code, this function randomly shuffles your dataset before splitting it into training and test sets. That means:
- Your
X_train/y_trainandX_test/y_testare different every run - Since KNN relies entirely on the training data to find nearest neighbors, changing the training set changes which neighbors get picked for predictions
- This leads to different prediction results, and thus different accuracy scores
Quick Fix: Fix the Random Seed
To make your splits (and therefore your accuracy) consistent every time you run the code, just add the random_state parameter to train_test_split. Pick any integer (like 42, a common choice for reproducibility) and set it here—this locks in the random shuffle pattern.
Corrected Code
I also fixed a couple of small issues in your original code (like missing the import for train_test_split and unused imports):
# Removed unused imports (scipy, numpy) since you aren't using them here from sklearn import datasets from sklearn.model_selection import train_test_split # Added missing import from sklearn.neighbors import KNeighborsClassifier from sklearn.metrics import accuracy_score iris = datasets.load_iris() X = iris.data y = iris.target # Added random_state to fix the split X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.5, random_state=42) my_classifier = KNeighborsClassifier() my_classifier.fit(X_train, y_train) predictions = my_classifier.predict(X_test) print(predictions) print(accuracy_score(y_test, predictions))
Now every time you run this, your training/test sets will be identical, so your accuracy score will stay the same.
Bonus Note
If you want to make sure your model's performance is robust (not just tied to one specific train/test split), you can use techniques like cross-validation (e.g., cross_val_score from sklearn) to average accuracy across multiple splits. But for reproducibility in single runs, fixing random_state is the way to go.
内容的提问来源于stack exchange,提问作者Ashish.k

