如何从Keras的model或hist中提取训练集与验证集的样本索引?
Hey there! Let's walk through how to figure out which samples from your original X and Y went to training vs validation. The approach depends entirely on how you set up your training workflow, since Keras doesn't store these indices in the model or History object by default.
Scenario 1: You used validation_split in model.fit()
When you use validation_split=0.2 (for example), Keras splits your data into training and validation sets automatically:
- If
shuffle=False(default is True for training), Keras takes the last 20% of your data as validation. You can calculate indices directly:import numpy as np n_samples = len(X) val_split = 0.2 train_indices = np.arange(int(n_samples * (1 - val_split))) val_indices = np.arange(int(n_samples * (1 - val_split)), n_samples) - If
shuffle=True(the default), Keras shuffles your data first before splitting. This means you can't recover the original indices from the model or History object—you'll need to manually shuffle and track indices before training:import numpy as np # Shuffle indices first indices = np.random.permutation(len(X)) split_idx = int(len(X) * (1 - val_split)) train_indices = indices[:split_idx] val_indices = indices[split_idx:] # Use these indices to get your split data X_train, X_val = X[train_indices], X[val_indices] y_train, y_val = y[train_indices], y[val_indices] # Train with the pre-split data (no need for validation_split) hist = model.fit(X_train, y_train, validation_data=(X_val, y_val), ...)
Scenario 2: You manually split your data (e.g., with train_test_split)
If you used a tool like sklearn.model_selection.train_test_split to split your data before training, capture the indices directly during the split:
from sklearn.model_selection import train_test_split # Split data AND capture original indices X_train, X_val, y_train, y_val, train_indices, val_indices = train_test_split( X, y, np.arange(len(X)), test_size=0.2, random_state=42 )
Here, train_indices and val_indices are exactly the indices from your original X/Y that went to training and validation, respectively. You already have what you need—no need to dig into the model or History object.
Scenario 3: You're using data generators (e.g., ImageDataGenerator)
If you're using generators like flow_from_directory or flow:
- For
flow, passshuffle=Falseto preserve the order of your input data, or attach original indices to your data and extract them from the generator output. - For
flow_from_directory, use the generator'sfilenamesattribute to get the list of files in the order they're processed. Map these filenames back to your original dataset's indices if you have a record of which file corresponds to which index.
Important Note
Keras' History object (your hist variable) only stores training metrics like loss and accuracy over epochs—it doesn't track sample indices. The model object itself also doesn't store this information, since it only processes the data you pass during training.
内容的提问来源于stack exchange,提问作者yohai

