You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拼接不同尺寸的NumPy数组?解决X_train与y_train拼接洗牌需求

How to Shuffle X_train and y_train Together

Hey there! Let's work through how to shuffle your training data properly. The main thing to note here is the shape of your X_train—it's currently (3072, 50000) (features first, samples second), while standard ML data formats usually have samples first. That's probably why your previous attempts didn't work. Let's cover two approaches:

You don't actually need to concatenate the arrays to shuffle them together. This is the better approach because it keeps your features and labels separate (avoiding unnecessary type conversions) and is more efficient.

Here's how to do it with NumPy:

import numpy as np

# First, transpose X_train to get samples first: (50000, 3072)
X_train = X_train.T

# Generate a shuffled list of indices for your samples
shuffled_indices = np.random.permutation(X_train.shape[0])

# Apply the indices to both X and y to shuffle them in sync
X_train_shuffled = X_train[shuffled_indices]
y_train_shuffled = y_train[shuffled_indices]

This works because we're using the same random order of indices for both arrays—so each feature vector in X stays paired with its corresponding label in y.

Method 2: Concatenate First (If You Really Need To)

If you must combine the arrays into a single structure, you'll need to adjust their shapes to match first:

import numpy as np

# Transpose X_train to (50000, 3072)
X_train = X_train.T

# Reshape y_train from (50000,) to (50000, 1) so it has the same number of rows as X
y_train_2d = y_train.reshape(-1, 1)

# Concatenate along the column axis
combined_data = np.hstack((X_train, y_train_2d))

# Shuffle the combined array in place
np.random.shuffle(combined_data)

# Optional: Split back into X and y if needed
X_train_shuffled = combined_data[:, :-1]
y_train_shuffled = combined_data[:, -1].flatten()

Why Your Previous Attempts Failed

  • Wrong shape alignment: Your original X_train has features as the first dimension, while y_train uses samples as the first dimension—they can't be concatenated directly without transposing X_train.
  • 1D vs 2D mismatch: np.hstack requires arrays to have the same number of rows. Since y_train was 1D, you need to reshape it to 2D to match the transposed X_train.

内容的提问来源于stack exchange,提问作者Monty _s Flying Circus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:02:03