Python从零实现K折交叉验证拆分及拆分标记、NaN准确率问题求助
Hey there! Let's break down your questions with practical code and clear explanations:
If you want to build K-Fold splitting without relying on libraries like sklearn, here's a straightforward, customizable approach. The core idea is to shuffle your data (to avoid order bias), split it into K roughly equal parts, then iterate to use each part as the validation set while the rest act as training data.
Step-by-Step Code Implementation
import numpy as np def custom_kfold_split(X, y, k=5, shuffle=True, random_state=42): n_samples = X.shape[0] indices = np.arange(n_samples) # Shuffle indices to prevent order-based bias (e.g., sorted labels) if shuffle: np.random.seed(random_state) np.random.shuffle(indices) fold_size = n_samples // k remainder = n_samples % k # Handle cases where sample count isn't divisible by k folds = [] start_idx = 0 for fold in range(k): # Assign extra samples to the first 'remainder' folds for balanced splits end_idx = start_idx + fold_size + (1 if fold < remainder else 0) val_indices = indices[start_idx:end_idx] train_indices = np.concatenate([indices[:start_idx], indices[end_idx:]]) folds.append((train_indices, val_indices)) start_idx = end_idx return folds # Example usage with dummy data (replace with your actual dataset) X = np.random.rand(100, 10) # 100 samples, 10 features y = np.random.randint(0, 2, size=100) # Binary classification labels # Generate 5-fold splits kfolds = custom_kfold_split(X, y, k=5) # Iterate through each fold to verify splits for fold_num, (train_idx, val_idx) in enumerate(kfolds, 1): X_train, X_val = X[train_idx], X[val_idx] y_train, y_val = y[train_idx], y[val_idx] print(f"Fold {fold_num}: Training samples = {len(X_train)}, Validation samples = {len(X_val)}")
Key Notes
- Shuffling is optional but highly recommended to ensure your splits don't inherit any hidden order patterns in your data.
- The
remaindervariable ensures splits stay balanced even when your total sample count isn't a perfect multiple of K.
How to Mark/Identify Each Fold Split
Tracking each fold is simple with a counter variable, and you can store fold-specific data/results in a dictionary or list for easy access later:
# Store fold data in a list for quick reference fold_data = [] for fold_num, (train_idx, val_idx) in enumerate(kfolds, 1): X_train, X_val = X[train_idx], X[val_idx] y_train, y_val = y[train_idx], y[val_idx] fold_data.append({ "fold_number": fold_num, "X_train": X_train, "X_val": X_val, "y_train": y_train, "y_val": y_val }) # Example: Access data for Fold 3 fold_3 = next(fold for fold in fold_data if fold["fold_number"] == 3)
When training your model, log results per fold to keep track of performance across splits:
from sklearn.linear_model import LogisticRegression from sklearn.metrics import accuracy_score from sklearn.preprocessing import StandardScaler fold_results = {} for fold in fold_data: fold_num = fold["fold_number"] X_train, X_val = fold["X_train"], fold["X_val"] y_train, y_val = fold["y_train"], fold["y_val"] # Critical preprocessing: Standardize features for logistic regression scaler = StandardScaler() X_train_scaled = scaler.fit_transform(X_train) X_val_scaled = scaler.transform(X_val) # Train model with increased max_iter to avoid convergence issues model = LogisticRegression(max_iter=1000, random_state=42) model.fit(X_train_scaled, y_train) # Predict and calculate accuracy y_pred = model.predict(X_val_scaled) acc = accuracy_score(y_val, y_pred) fold_results[f"Fold {fold_num}"] = round(acc, 4) print(f"Fold {fold_num} Accuracy: {acc:.4f}") print("\nFinal Fold Results:", fold_results)
Fixing NaN Accuracy
NaN accuracy almost always comes from one of these issues—here's how to troubleshoot:
Issue 1: Missing Values in Data
If your features or labels contain NaNs, the model may output NaN predictions. Check for missing values with:# For numpy arrays print("Missing values in X:", np.isnan(X).sum()) print("Missing values in y:", np.isnan(y).sum())Fix this by imputing missing values (e.g., using
SimpleImputerfrom sklearn) or removing samples with missing data.Issue 2: Model Failed to Converge
Logistic regression needs enough iterations to converge. Without standardized features or a lowmax_iter, the model may not converge, leading to NaN parameters/predictions:- Always standardize features for logistic regression (it speeds up convergence).
- Increase
max_iter(e.g., to 1000 or higher) and check convergence withmodel.n_iter_.
Issue 3: Bug in Accuracy Calculation
Double-check your code—did you mix upy_valandy_pred? Are you using the wrong metric? For example,roc_auc_scorereturns NaN for single-class data, butaccuracy_scorewon't.Issue 4: Single-Class Validation Set
If a validation set has only one class (e.g., all 0s),accuracy_scorewill return 1.0 or 0.0 (not NaN). If you're seeing NaN here, it's likely tied to the first two issues above.
内容的提问来源于stack exchange,提问作者DN1

