如何创建新数据帧?解决值(1000,10)与索引(1000,11)维度不匹配问题
Hey there! Let's break down your two questions clearly—first, how to create a new pandas DataFrame, then fixing that frustrating shape mismatch error you're running into.
There are a few common, straightforward ways to make a new DataFrame, depending on your source data:
- From scratch using dictionaries: Perfect for small datasets you want to define manually.
import pandas as pd data = { 'Name': ['Alice', 'Bob', 'Charlie'], 'Age': [25, 30, 35], 'City': ['New York', 'London', 'Paris'] } new_df = pd.DataFrame(data) - From NumPy arrays or other tabular data: If you have a numerical array (like your scaled features), just specify matching column names when creating the DataFrame.
- From an existing DataFrame: Slice or filter the original data to get a subset, and use
.copy()to avoid accidentally modifying the original dataset:# Create a new DataFrame with specific columns from the original subset_df = original_df[['Column1', 'Column3']].copy()
Let's start with why this error pops up:
You're scaling features after dropping the TARGET CLASS column (leaving 10 columns total), but then you're trying to assign all 11 columns from the original DataFrame (df.columns) to the scaled data. Pandas throws an error because the number of columns doesn't line up.
Here's the corrected code, with comments explaining the fixes:
import pandas as pd # Don't forget to import pandas if you haven't already! from sklearn.preprocessing import StandardScaler scaler = StandardScaler() # Store the feature columns (excluding TARGET CLASS) once to avoid repetition feature_columns = df.drop('TARGET CLASS', axis=1).columns # Fit the scaler on just the feature columns scaler.fit(df[feature_columns]) # Transform the same feature columns scaled_features = scaler.transform(df[feature_columns]) # Now create the DataFrame with the correct feature column names (10 columns, not 11) df_feat = pd.DataFrame(scaled_features, columns=feature_columns)
Optional: Add the Target Column Back
If you want to include the TARGET CLASS in your new DataFrame, you can append it after creating df_feat:
df_feat['TARGET CLASS'] = df['TARGET CLASS'].values
That should resolve the shape mismatch error entirely—you're now matching the number of columns in your scaled data to the column names you're assigning.
内容的提问来源于stack exchange,提问作者Didarbek Nuraliyev

