Python按组构建加权线性回归遇样本权重维度错误求助
问题
现有一个基于DataFrame的可正常运行的回归模型,自变量与因变量均来自DataFrame。需要为模型添加百分比格式的权重,已在DataFrame中新增Weight列存储权重值。要求按sc列分组,为每个分组创建带对应权重的回归模型及方程,但运行时触发错误。
相关代码
import pandas as pd from sklearn.linear_model import LinearRegression df = pd.read_excel('Weighted Test.xlsx') grouped = df.groupby('sc') sc = [] region = [] equations = [] m_value = [] c_value = [] df['Avg_actual_cube'].fillna(0, inplace=True) df['Actual_TPH'].fillna(0, inplace=True) df['Weight'].fillna(0, inplace=True) grouped.head()
数据样例输出
sc Region Week of balance_day Year of balance_day Actual_TPH Avg_actual_cube Weight 0 ABQ5 West Week 53 2023 2023 122.443157 0.490000 0.1 1 ABQ5 West Week 36 2023 2023 96.074478 0.445714 0.2 2 ABQ5 West Week 22 2024 2024 87.164843 0.510000 0.3
加权回归代码
sample_weight=df['Weight'] for group_name, group_data in grouped: X = group_data['Avg_actual_cube'].values.reshape(-1, 1) y = group_data['Actual_TPH'].values weighted_model = LinearRegression() model.fit(X, y,sample_weight) sc.append(group_name) region.append(group_data['Region'].values[0]) equations.append('TPH = {:.2f}cube+ {:.2f}'.format(model.coef_[0], model.intercept_)) m_value.append(model.coef_[0]) c_value.append(model.intercept_)
触发的错误
ValueError Traceback (most recent call last) Input In [14], in <cell line: 1>() 3 y = group_data['Actual_TPH'].values 4 weighted_model = LinearRegression() ----> 5 model.fit(X, y,sample_weight) 7 sc.append(group_name) 8 region.append(group_data['Region'].values[0]) File ~\Anaconda3\lib\site-packages\sklearn\linear_model\_base.py:667, in LinearRegression.fit(self, X, y, sample_weight) 662 X, y = self._validate_data( 663 X, y, accept_sparse=accept_sparse, y_numeric=True, multi_output=True 664 ) 666 if sample_weight is not None: --> 667 sample_weight = _check_sample_weight(sample_weight, X, dtype=X.dtype) 669 X, y, X_offset, y_offset, X_scale = self._preprocess_data( 670 X, 671 y, (...) 676 return_mean=True, 677 ) 679 if sample_weight is not None: 680 # Sample weight can be implemented via a simple rescaling. File ~\Anaconda3\lib\site-packages\sklearn\utils\validation.py:1565, in _check_sample_weight(sample_weight, X, dtype, copy) 1562 raise ValueError("Sample weights must be 1D array or scalar") 1564 if sample_weight.shape != (n_samples,): --> 1565 raise ValueError( 1566 "sample_weight.shape == {}, expected {}!".format( 1567 sample_weight.shape, (n_samples,) 1568 ) 1569 ) 1571 return sample_weight ValueError: sample_weight.shape == (5967,), expected (52,)!
解决方案
问题原因
错误核心是权重维度不匹配:
- 你定义的
sample_weight=df['Weight']是整个DataFrame的权重列,长度为5967,对应全量数据的行数。 - 但循环中每个分组的样本量只有52行,模型要求权重的长度必须和当前分组的样本数完全一致,因此触发维度不匹配错误。
- 同时存在变量名错误:实例化了
weighted_model,但调用的是未定义的model.fit()。
修正后的代码
# 移除全局的sample_weight定义,改为在循环内获取对应分组的权重 for group_name, group_data in grouped: X = group_data['Avg_actual_cube'].values.reshape(-1, 1) y = group_data['Actual_TPH'].values # 关键:从当前分组中提取对应的权重 sample_weight = group_data['Weight'].values weighted_model = LinearRegression() # 使用当前分组的权重拟合模型,修正变量名错误 weighted_model.fit(X, y, sample_weight) sc.append(group_name) region.append(group_data['Region'].values[0]) equations.append('TPH = {:.2f}cube + {:.2f}'.format(weighted_model.coef_[0], weighted_model.intercept_)) m_value.append(weighted_model.coef_[0]) c_value.append(weighted_model.intercept_)
额外说明
- 你已经用
fillna(0)处理了Weight列的缺失值,这部分无需调整。 - 修正后每个分组的权重长度会和当前分组的样本数完全匹配,满足模型的输入要求。
内容的提问来源于stack exchange,提问作者Gaurav Hegde
相关产品推荐
相关产品推荐

