在DataFrame上执行OLS回归时遇ValueError:维度不匹配求助
Got it, let's break down this dimension mismatch issue you're hitting. That error message (shapes (48,34) and (48,34) not aligned: 34 (dim 1) != 48 (dim 0)) tells us one key thing: one or more of your independent variables (ownership, shipping, title) isn't a single-column Series—it's actually a 34-column DataFrame. Statsmodels expects each variable in your formula to be 1-dimensional (one value per row), but when it encounters a multi-column variable, it tries to include all those columns in the design matrix, leading to the misalignment during matrix multiplication.
Step 1: Diagnose the problematic variable
First, let's confirm which variable is causing the issue. Run these quick checks on your sold1 DataFrame:
# Check overall DataFrame structure and column types print(sold1.info()) # Check the shape of each individual column for col in sold1.columns: print(f"Column '{col}' shape: {sold1[col].shape}")
Look for any column that outputs a shape like (48, 34)—that's the culprit. Common reasons this happens:
- You accidentally stored a multi-column result (like one-hot encoded data) under an existing column name
- A column was parsed as a nested DataFrame when loading your data
- You have duplicate column names (statsmodels will combine all columns with the same name into a multi-column variable)
Step 2: Fix the multi-column variable
Once you've identified the problematic column, fix it based on the root cause:
- If it's a nested DataFrame: Extract the single column you actually want to use. For example, if
ownershipis the multi-column variable and you need the first column:# Replace the multi-column variable with a single column sold1['ownership'] = sold1['ownership'].iloc[:, 0] - If there are duplicate columns: Rename the duplicates to unique names, then adjust your formula to use the correct one:
# Check for duplicate column names print(sold1.columns.duplicated()) # Rename duplicates to unique names sold1.columns = [f"{col}_{i}" if dup else col for i, (col, dup) in enumerate(zip(sold1.columns, sold1.columns.duplicated()))] - If it's unintended multi-column data: Revert the column to its original single-column state (e.g., undo a one-hot encoding step if you didn't mean to apply it to that variable).
Step 3: Re-run your OLS model
After ensuring all variables in your formula are single-column Series, re-run your code (note: we usually use smf as the alias for statsmodels.formula.api for clarity):
import numpy as np import statsmodels.formula.api as smf result = smf.ols(formula="price ~ ownership + shipping + title", data=sold1).fit() print(result.summary())
That should resolve the dimension alignment error and let you fit the model successfully.
内容的提问来源于stack exchange,提问作者Nazmul Islam

