You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按组构建加权线性回归遇样本权重维度错误求助

问题

现有一个基于DataFrame的可正常运行的回归模型,自变量与因变量均来自DataFrame。需要为模型添加百分比格式的权重,已在DataFrame中新增Weight列存储权重值。要求按sc列分组,为每个分组创建带对应权重的回归模型及方程,但运行时触发错误。

相关代码

import pandas as pd
from sklearn.linear_model import LinearRegression

df = pd.read_excel('Weighted Test.xlsx')
grouped = df.groupby('sc')
sc = []
region = []
equations = []
m_value = []
c_value = []
df['Avg_actual_cube'].fillna(0, inplace=True)
df['Actual_TPH'].fillna(0, inplace=True)
df['Weight'].fillna(0, inplace=True)
grouped.head()

数据样例输出

sc  Region  Week of balance_day Year of balance_day  Actual_TPH  Avg_actual_cube  Weight
0  ABQ5    West       Week 53 2023                2023  122.443157          0.490000     0.1
1  ABQ5    West       Week 36 2023                2023   96.074478          0.445714     0.2
2  ABQ5    West       Week 22 2024                2024   87.164843          0.510000     0.3

加权回归代码

sample_weight=df['Weight']

for group_name, group_data in grouped:
    X = group_data['Avg_actual_cube'].values.reshape(-1, 1)
    y = group_data['Actual_TPH'].values
    weighted_model = LinearRegression()
    model.fit(X, y,sample_weight)
    sc.append(group_name)
    region.append(group_data['Region'].values[0])
    equations.append('TPH = {:.2f}cube+ {:.2f}'.format(model.coef_[0], model.intercept_))
    m_value.append(model.coef_[0])
    c_value.append(model.intercept_)

触发的错误

ValueError                                Traceback (most recent call last)
Input In [14], in <cell line: 1>()
      3 y = group_data['Actual_TPH'].values
      4 weighted_model = LinearRegression()
----> 5 model.fit(X, y,sample_weight)
      7 sc.append(group_name)
      8 region.append(group_data['Region'].values[0])

File ~\Anaconda3\lib\site-packages\sklearn\linear_model\_base.py:667, in LinearRegression.fit(self, X, y, sample_weight)
    662 X, y = self._validate_data(
    663     X, y, accept_sparse=accept_sparse, y_numeric=True, multi_output=True
    664 )
    666 if sample_weight is not None:
--> 667     sample_weight = _check_sample_weight(sample_weight, X, dtype=X.dtype)
    669 X, y, X_offset, y_offset, X_scale = self._preprocess_data(
    670     X,
    671     y,
   (...)
    676     return_mean=True,
    677 )
    679 if sample_weight is not None:
    680     # Sample weight can be implemented via a simple rescaling.

File ~\Anaconda3\lib\site-packages\sklearn\utils\validation.py:1565, in _check_sample_weight(sample_weight, X, dtype, copy)
   1562         raise ValueError("Sample weights must be 1D array or scalar")
   1564     if sample_weight.shape != (n_samples,):
--> 1565         raise ValueError(
   1566             "sample_weight.shape == {}, expected {}!".format(
   1567                 sample_weight.shape, (n_samples,)
   1568             )
   1569         )
   1571 return sample_weight

ValueError: sample_weight.shape == (5967,), expected (52,)!
解决方案

问题原因

错误核心是权重维度不匹配:

  • 你定义的sample_weight=df['Weight']是整个DataFrame的权重列,长度为5967,对应全量数据的行数。
  • 但循环中每个分组的样本量只有52行,模型要求权重的长度必须和当前分组的样本数完全一致,因此触发维度不匹配错误。
  • 同时存在变量名错误:实例化了weighted_model,但调用的是未定义的model.fit()。

修正后的代码

# 移除全局的sample_weight定义,改为在循环内获取对应分组的权重
for group_name, group_data in grouped:
    X = group_data['Avg_actual_cube'].values.reshape(-1, 1)
    y = group_data['Actual_TPH'].values
    # 关键:从当前分组中提取对应的权重
    sample_weight = group_data['Weight'].values
    weighted_model = LinearRegression()
    # 使用当前分组的权重拟合模型,修正变量名错误
    weighted_model.fit(X, y, sample_weight)
    sc.append(group_name)
    region.append(group_data['Region'].values[0])
    equations.append('TPH = {:.2f}cube + {:.2f}'.format(weighted_model.coef_[0], weighted_model.intercept_))
    m_value.append(weighted_model.coef_[0])
    c_value.append(weighted_model.intercept_)

额外说明

  • 你已经用fillna(0)处理了Weight列的缺失值,这部分无需调整。
  • 修正后每个分组的权重长度会和当前分组的样本数完全匹配,满足模型的输入要求。

内容的提问来源于stack exchange,提问作者Gaurav Hegde

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 00:32:07