You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从零实现K折交叉验证拆分及拆分标记、NaN准确率问题求助

Hey there! Let's break down your questions with practical code and clear explanations:

1. Implementing K-Fold Cross Validation from Scratch in Python

If you want to build K-Fold splitting without relying on libraries like sklearn, here's a straightforward, customizable approach. The core idea is to shuffle your data (to avoid order bias), split it into K roughly equal parts, then iterate to use each part as the validation set while the rest act as training data.

Step-by-Step Code Implementation

import numpy as np

def custom_kfold_split(X, y, k=5, shuffle=True, random_state=42):
    n_samples = X.shape[0]
    indices = np.arange(n_samples)
    
    # Shuffle indices to prevent order-based bias (e.g., sorted labels)
    if shuffle:
        np.random.seed(random_state)
        np.random.shuffle(indices)
    
    fold_size = n_samples // k
    remainder = n_samples % k  # Handle cases where sample count isn't divisible by k
    folds = []
    start_idx = 0
    
    for fold in range(k):
        # Assign extra samples to the first 'remainder' folds for balanced splits
        end_idx = start_idx + fold_size + (1 if fold < remainder else 0)
        val_indices = indices[start_idx:end_idx]
        train_indices = np.concatenate([indices[:start_idx], indices[end_idx:]])
        
        folds.append((train_indices, val_indices))
        start_idx = end_idx
    
    return folds

# Example usage with dummy data (replace with your actual dataset)
X = np.random.rand(100, 10)  # 100 samples, 10 features
y = np.random.randint(0, 2, size=100)  # Binary classification labels

# Generate 5-fold splits
kfolds = custom_kfold_split(X, y, k=5)

# Iterate through each fold to verify splits
for fold_num, (train_idx, val_idx) in enumerate(kfolds, 1):
    X_train, X_val = X[train_idx], X[val_idx]
    y_train, y_val = y[train_idx], y[val_idx]
    print(f"Fold {fold_num}: Training samples = {len(X_train)}, Validation samples = {len(X_val)}")

Key Notes

  • Shuffling is optional but highly recommended to ensure your splits don't inherit any hidden order patterns in your data.
  • The remainder variable ensures splits stay balanced even when your total sample count isn't a perfect multiple of K.

2. Identifying Fold Splits & Fixing NaN Accuracy

How to Mark/Identify Each Fold Split

Tracking each fold is simple with a counter variable, and you can store fold-specific data/results in a dictionary or list for easy access later:

# Store fold data in a list for quick reference
fold_data = []
for fold_num, (train_idx, val_idx) in enumerate(kfolds, 1):
    X_train, X_val = X[train_idx], X[val_idx]
    y_train, y_val = y[train_idx], y[val_idx]
    
    fold_data.append({
        "fold_number": fold_num,
        "X_train": X_train,
        "X_val": X_val,
        "y_train": y_train,
        "y_val": y_val
    })

# Example: Access data for Fold 3
fold_3 = next(fold for fold in fold_data if fold["fold_number"] == 3)

When training your model, log results per fold to keep track of performance across splits:

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.preprocessing import StandardScaler

fold_results = {}

for fold in fold_data:
    fold_num = fold["fold_number"]
    X_train, X_val = fold["X_train"], fold["X_val"]
    y_train, y_val = fold["y_train"], fold["y_val"]
    
    # Critical preprocessing: Standardize features for logistic regression
    scaler = StandardScaler()
    X_train_scaled = scaler.fit_transform(X_train)
    X_val_scaled = scaler.transform(X_val)
    
    # Train model with increased max_iter to avoid convergence issues
    model = LogisticRegression(max_iter=1000, random_state=42)
    model.fit(X_train_scaled, y_train)
    
    # Predict and calculate accuracy
    y_pred = model.predict(X_val_scaled)
    acc = accuracy_score(y_val, y_pred)
    
    fold_results[f"Fold {fold_num}"] = round(acc, 4)
    print(f"Fold {fold_num} Accuracy: {acc:.4f}")

print("\nFinal Fold Results:", fold_results)

Fixing NaN Accuracy

NaN accuracy almost always comes from one of these issues—here's how to troubleshoot:

  • Issue 1: Missing Values in Data
    If your features or labels contain NaNs, the model may output NaN predictions. Check for missing values with:

    # For numpy arrays
    print("Missing values in X:", np.isnan(X).sum())
    print("Missing values in y:", np.isnan(y).sum())
    

    Fix this by imputing missing values (e.g., using SimpleImputer from sklearn) or removing samples with missing data.

  • Issue 2: Model Failed to Converge
    Logistic regression needs enough iterations to converge. Without standardized features or a low max_iter, the model may not converge, leading to NaN parameters/predictions:

    • Always standardize features for logistic regression (it speeds up convergence).
    • Increase max_iter (e.g., to 1000 or higher) and check convergence with model.n_iter_.
  • Issue 3: Bug in Accuracy Calculation
    Double-check your code—did you mix up y_val and y_pred? Are you using the wrong metric? For example, roc_auc_score returns NaN for single-class data, but accuracy_score won't.

  • Issue 4: Single-Class Validation Set
    If a validation set has only one class (e.g., all 0s), accuracy_score will return 1.0 or 0.0 (not NaN). If you're seeing NaN here, it's likely tied to the first two issues above.


内容的提问来源于stack exchange,提问作者DN1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:26:38