Numpy Reshape维度转换异常:(8,820,5)转(820,8,5)后数据不符
Hey there, let's break down why you're seeing that unexpected behavior with reshape—it's all about how numpy orders elements when reshaping vs rearranging axes.
First, let's recap your original data structure:
- Initial
inputs.shape = (8, 820, 5): This means 8 stocks × 820 days × 5 features (with adj_close as the first feature). - When you extract
prices = inputs[:, :, 0], you get an array of shape(8,820): each row is a single stock's daily closing prices over 820 days. Soprices[:,0]correctly gives you all 8 stocks' day-0 close prices.
Why reshape(820,8,5) isn't working as expected
Numpy's reshape uses row-major (C-style) ordering by default. That means it first flattens your original array into a 1D sequence, then packs those elements into the new shape. Let's see what that looks like for your data:
Original element order (flattened):[stock0_day0_feat0, stock0_day0_feat1, ..., stock0_day0_feat4, stock0_day1_feat0, ..., stock0_day819_feat4, stock1_day0_feat0, ..., stock7_day819_feat4]
When you reshape to (820,8,5), numpy splits this flattened list into 820 groups, each with 8×5 elements. The first group (index 0 in the new array) will be:[stock0_day0_feat0, stock0_day0_feat1, ..., stock0_day0_feat4, stock0_day1_feat0, ..., stock0_day1_feat4, ..., stock0_day7_feat4]
So when you extract prices = inputs[:, :, 0] after reshaping, prices[0] ends up being stock0's first 8 days of close prices—not all stocks' day-0 prices. That's why you're seeing the unexpected result!
The fix: Use transpose instead of reshape
What you actually want is to swap the axes so that days come first, then stocks, then features. Numpy's transpose lets you reorder the axes directly without flattening the array.
Your original axes are:
- Axis 0: Stocks (8)
- Axis 1: Days (820)
- Axis 2: Features (5)
To get the desired shape (820,8,5), you need to reorder the axes to Days → Stocks → Features. That means swapping axis 0 and axis 1. Here's how to do it:
# Original shape: (8, 820, 5) inputs_transposed = inputs.transpose(1, 0, 2) # Now inputs_transposed.shape = (820, 8, 5)
Now, when you extract prices:
prices = inputs_transposed[:, :, 0] # prices.shape = (820, 8) prices[0] # This gives you all 8 stocks' day-0 closing prices—exactly what you want!
Why this works
transpose(1,0,2) maps each element from the original position (stock_idx, day_idx, feat_idx) to the new position (day_idx, stock_idx, feat_idx). So every entry for day d across all stocks is grouped together in the first dimension of the transposed array.
Quick example to visualize
Let's use a tiny dataset to make this concrete:
- Original shape:
(2 stocks, 3 days, 1 feature)→[[[1], [2], [3]], [[4], [5], [6]]] - Flattened order:
[1,2,3,4,5,6]
If we reshape to (3,2,1):
reshaped = original.reshape(3,2,1) # Result: [[[1], [2]], [[3], [4]], [[5], [6]]] # Each row is consecutive days from the first stock, then the second
If we transpose to (3,2,1) (axes 1,0,2):
transposed = original.transpose(1,0,2) # Result: [[[1], [4]], [[2], [5]], [[3], [6]]] # Each row is all stocks' data for a single day—perfect for your use case!
Calculating log-returns
Now that you have the correct price array shape (820,8), calculating log-returns is straightforward. For each stock (column), compute the log of the ratio between consecutive days:
import numpy as np log_returns = np.log(prices[1:] / prices[:-1]) # log_returns.shape = (819, 8) → 819 daily returns for each of the 8 stocks
That should give you the log-returns you need for your adj_close data.
内容的提问来源于stack exchange,提问作者Daniel Chepenko

