You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在DataFrame上执行OLS回归时遇ValueError:维度不匹配求助

Fixing the "shapes not aligned" ValueError in statsmodels OLS

Got it, let's break down this dimension mismatch issue you're hitting. That error message (shapes (48,34) and (48,34) not aligned: 34 (dim 1) != 48 (dim 0)) tells us one key thing: one or more of your independent variables (ownership, shipping, title) isn't a single-column Series—it's actually a 34-column DataFrame. Statsmodels expects each variable in your formula to be 1-dimensional (one value per row), but when it encounters a multi-column variable, it tries to include all those columns in the design matrix, leading to the misalignment during matrix multiplication.

Step 1: Diagnose the problematic variable

First, let's confirm which variable is causing the issue. Run these quick checks on your sold1 DataFrame:

# Check overall DataFrame structure and column types
print(sold1.info())

# Check the shape of each individual column
for col in sold1.columns:
    print(f"Column '{col}' shape: {sold1[col].shape}")

Look for any column that outputs a shape like (48, 34)—that's the culprit. Common reasons this happens:

  • You accidentally stored a multi-column result (like one-hot encoded data) under an existing column name
  • A column was parsed as a nested DataFrame when loading your data
  • You have duplicate column names (statsmodels will combine all columns with the same name into a multi-column variable)

Step 2: Fix the multi-column variable

Once you've identified the problematic column, fix it based on the root cause:

  • If it's a nested DataFrame: Extract the single column you actually want to use. For example, if ownership is the multi-column variable and you need the first column:
    # Replace the multi-column variable with a single column
    sold1['ownership'] = sold1['ownership'].iloc[:, 0]
    
  • If there are duplicate columns: Rename the duplicates to unique names, then adjust your formula to use the correct one:
    # Check for duplicate column names
    print(sold1.columns.duplicated())
    # Rename duplicates to unique names
    sold1.columns = [f"{col}_{i}" if dup else col for i, (col, dup) in enumerate(zip(sold1.columns, sold1.columns.duplicated()))]
    
  • If it's unintended multi-column data: Revert the column to its original single-column state (e.g., undo a one-hot encoding step if you didn't mean to apply it to that variable).

Step 3: Re-run your OLS model

After ensuring all variables in your formula are single-column Series, re-run your code (note: we usually use smf as the alias for statsmodels.formula.api for clarity):

import numpy as np
import statsmodels.formula.api as smf

result = smf.ols(formula="price ~ ownership + shipping + title", data=sold1).fit()
print(result.summary())

That should resolve the dimension alignment error and let you fit the model successfully.

内容的提问来源于stack exchange,提问作者Nazmul Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:52:51