如何将含多维数组的Pandas Series转换为DataFrame或指定维度Numpy数组
Hey there! Let's break this down for you since you're new to Python and Pandas. You've got a Series where each entry is a (18,100) array, and converting it directly with pd.DataFrame(pds) just shoves each array into a single column—here's how to get the 3D NumPy array or a more useful DataFrame you're looking for:
Since every element in your Series has the exact same shape ((18,100)), NumPy's stack() function is perfect here. It'll take all those 2D arrays and stack them along a new first dimension to make a 3D array:
import numpy as np # Convert the Series values to a 3D NumPy array arr_3d = np.stack(pds.values) # Check the shape—this should return (1159, 18, 100) print(arr_3d.shape)
How you structure the DataFrame depends on what you need to do with the data. Here are two common options:
Option 1: Flatten the arrays into a 2D DataFrame
If you want each (18,100) array to be a single row with 1800 columns (18*100), you can reshape the 3D array we just made into a 2D structure, then wrap it in a DataFrame:
import pandas as pd # Flatten the 3D array to 2D (1159 rows × 1800 columns) df_flat = pd.DataFrame(arr_3d.reshape(len(pds), -1), index=pds.index) # Verify the shape print(df_flat.shape) # Output: (1159, 1800)
This keeps your original Series index, so you can still map rows back to the original entries.
Option 2: Use MultiIndex columns to preserve the array structure
If you want to keep track of the original (18,100) dimensions in your DataFrame, you can use a MultiIndex for columns. This makes it easier to slice data by the array's rows or columns later:
# Start with the 3D array we created arr_3d = np.stack(pds.values) # Create a DataFrame with flattened columns df_multi = pd.DataFrame(arr_3d.reshape(len(pds), -1), index=pds.index) # Set up a MultiIndex to reflect the original (18,100) shape df_multi.columns = pd.MultiIndex.from_product( [range(18), range(100)], names=["array_row", "array_col"] ) # Now you can slice by array rows or columns easily, e.g.: # Get all columns from the 5th row of the original arrays subset = df_multi.loc[:, (4, :)]
To quickly recap why pd.DataFrame(pds) gave you a (1159,1) shape: Pandas treats each array in your Series as a single scalar value, so it just puts each one into a single column. By stacking the arrays first, we give Pandas (or NumPy) the structured data it needs to build the shape you want.
内容的提问来源于stack exchange,提问作者Greeshma Agasthya

