You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从含50类的Numpy训练数据集中每类随机选取5个样本构建子数据集

Hey there! Creating a balanced sub-dataset with exactly 5 samples per class is straightforward with NumPy (or Pandas, if you prefer a more concise approach). Let's walk through how to do this step by step.

Step 1: Define parameters and confirm data structure

First, make sure your NumPy arrays are ready, then set your sampling parameters and get the list of unique classes:

import numpy as np

# Your original dataset (example values for context)
x_train = np.random.rand(9000, 2048)  # Shape: (9000, 2048)
y_train = np.array([f"class_{i%50}" for i in range(9000)])  # Shape: (9000,), string labels

# Key parameters
num_samples_per_class = 5
classes = list(set(y_train))  # List of your 50 unique classes

Step 2: Randomly select indices for each class

Loop through each class, find all indices of samples belonging to that class, then randomly pick 5 unique indices (no duplicates):

selected_indices = []

for cls in classes:
    # Get all indices where the label matches the current class
    class_indices = np.where(y_train == cls)[0]
    # Randomly select 5 indices without replacement
    sampled_indices = np.random.choice(class_indices, size=num_samples_per_class, replace=False)
    selected_indices.extend(sampled_indices)

# Convert the list to a NumPy array for easier indexing
selected_indices = np.array(selected_indices)

Step 3: Extract your sub-dataset

Use the selected indices to slice your original feature and label arrays:

sub_train_data = x_train[selected_indices]  # Shape: (250, 2048)
sub_train_labels = y_train[selected_indices]  # Shape: (250,)

Step 4: Verify the result

Double-check that each class has exactly 5 samples to ensure everything worked as expected:

unique_classes, sample_counts = np.unique(sub_train_labels, return_counts=True)
print(dict(zip(unique_classes, sample_counts)))
# Output should show each class with a count of 5

Bonus: Concise approach with Pandas

If you're comfortable using Pandas, this can be done in just a few lines with the groupby().sample() method:

import pandas as pd

# Convert NumPy arrays to a DataFrame
df = pd.DataFrame(x_train)
df["label"] = y_train

# Sample 5 rows per class
sub_df = df.groupby("label").sample(n=num_samples_per_class, random_state=42)

# Convert back to NumPy arrays
sub_train_data = sub_df.drop("label", axis=1).to_numpy()
sub_train_labels = sub_df["label"].to_numpy()

Quick Notes

  • Add np.random.seed(42) (or any integer) before selecting indices if you want your random sampling to be reproducible.
  • Using replace=False ensures you don't pick the same sample multiple times for a single class.

内容的提问来源于stack exchange,提问作者Joseph

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:19:00