如何从含50类的Numpy训练数据集中每类随机选取5个样本构建子数据集
Hey there! Creating a balanced sub-dataset with exactly 5 samples per class is straightforward with NumPy (or Pandas, if you prefer a more concise approach). Let's walk through how to do this step by step.
Step 1: Define parameters and confirm data structure
First, make sure your NumPy arrays are ready, then set your sampling parameters and get the list of unique classes:
import numpy as np # Your original dataset (example values for context) x_train = np.random.rand(9000, 2048) # Shape: (9000, 2048) y_train = np.array([f"class_{i%50}" for i in range(9000)]) # Shape: (9000,), string labels # Key parameters num_samples_per_class = 5 classes = list(set(y_train)) # List of your 50 unique classes
Step 2: Randomly select indices for each class
Loop through each class, find all indices of samples belonging to that class, then randomly pick 5 unique indices (no duplicates):
selected_indices = [] for cls in classes: # Get all indices where the label matches the current class class_indices = np.where(y_train == cls)[0] # Randomly select 5 indices without replacement sampled_indices = np.random.choice(class_indices, size=num_samples_per_class, replace=False) selected_indices.extend(sampled_indices) # Convert the list to a NumPy array for easier indexing selected_indices = np.array(selected_indices)
Step 3: Extract your sub-dataset
Use the selected indices to slice your original feature and label arrays:
sub_train_data = x_train[selected_indices] # Shape: (250, 2048) sub_train_labels = y_train[selected_indices] # Shape: (250,)
Step 4: Verify the result
Double-check that each class has exactly 5 samples to ensure everything worked as expected:
unique_classes, sample_counts = np.unique(sub_train_labels, return_counts=True) print(dict(zip(unique_classes, sample_counts))) # Output should show each class with a count of 5
Bonus: Concise approach with Pandas
If you're comfortable using Pandas, this can be done in just a few lines with the groupby().sample() method:
import pandas as pd # Convert NumPy arrays to a DataFrame df = pd.DataFrame(x_train) df["label"] = y_train # Sample 5 rows per class sub_df = df.groupby("label").sample(n=num_samples_per_class, random_state=42) # Convert back to NumPy arrays sub_train_data = sub_df.drop("label", axis=1).to_numpy() sub_train_labels = sub_df["label"].to_numpy()
Quick Notes
- Add
np.random.seed(42)(or any integer) before selecting indices if you want your random sampling to be reproducible. - Using
replace=Falseensures you don't pick the same sample multiple times for a single class.
内容的提问来源于stack exchange,提问作者Joseph

