多变量回归报错ValueError:样本数不一致问题排查求助
多变量回归样本数不匹配问题排查
问题背景
数据集来自300人对91件艺术品的调查:
- 前91列:1-7分的偏好评分
- 第92-182列:作品能量评分
- 第216-221列:人口统计数据
目标是基于能量评分和人口统计数据构建回归模型预测偏好评分,单变量回归可正常运行,但多变量回归执行时报错。
原代码
rate = data[:,0:91] #isolate preference ratings energy = data[:,91:182] #isolate energy ratings demo = data[:,215:221] #isolate demographic ratings demo[np.isnan(demo)] = 4.68 comb = np.concatenate((energy,demo),axis=1) x = comb.reshape(-1,1) y = rate.reshape(-1,1) ourModel = LinearRegression().fit(x, y) rSq = ourModel.score(x,y) slope = ourModel.coef_ intercept = ourModel.intercept_ yHat = slope * x + intercept print(rSq) plt.plot(x,y,'o',) plt.xlabel('Energy + Demographic') plt.ylabel('Preference') plt.plot(x,yHat,color='orange',linewidth=3) plt.title('Predicting Art Preference from Energy and Demographic') plt.show()
报错信息
ValueError: Found input variables with inconsistent numbers of samples: [29100, 27300]
报错原因
样本数不匹配的根源是数据重塑逻辑错误:
comb是将300行的energy(91列)和demo(6列)按列拼接,得到(300, 97)的数组,reshape(-1,1)后变成(300*97, 1) = (29100,1),这是错误的样本数。rate是(300,91)的数组,reshape(-1,1)后得到(300*91,1)=(27300,1),这是正确的样本数(每个用户对每件作品的评分对应一个样本)。
核心逻辑错误:没有建立每个偏好评分对应该作品能量评分+用户人口统计数据的正确映射,而是错误地把所有特征列强行压成单特征样本,导致样本数不匹配。
修正方案
需要将数据整理为:每个样本对应「1个偏好评分」+「该作品的能量评分+对应用户的人口统计数据」,具体步骤:
- 展开
rate为一维数组(27300个样本),保持每个评分的独立性。 - 展开
energy为一维数组(每个作品的能量评分对应到用户的评分样本)。 - 将
demo数据重复91次(每个用户的人口统计数据要对应他对91件作品的所有评分),得到和样本数匹配的(27300,6)数组。 - 拼接展开后的
energy和重复后的demo,得到特征矩阵X,形状为(27300,7),与y的样本数完全匹配。
修正后代码
import numpy as np from sklearn.linear_model import LinearRegression import matplotlib.pyplot as plt # 提取数据 rate = data[:,0:91] # 偏好评分 (300,91) energy = data[:,91:182] # 能量评分 (300,91) demo = data[:,215:221] # 人口统计数据 (300,6) # 填充缺失值 demo[np.isnan(demo)] = 4.68 # 处理特征和标签,建立正确的样本映射 y = rate.flatten() # 展开为(27300,)的一维数组,每个元素是一个偏好评分样本 energy_flat = energy.flatten() # 展开能量评分为(27300,) # 将人口统计数据重复91次,每个用户的demo对应91个作品评分 demo_repeated = np.repeat(demo, repeats=91, axis=0) # 形状变为(27300,6) # 拼接特征:能量评分(1列) + 人口统计(6列) X = np.column_stack((energy_flat, demo_repeated)) # 形状(27300,7) # 训练模型 ourModel = LinearRegression().fit(X, y) rSq = ourModel.score(X, y) slope = ourModel.coef_ intercept = ourModel.intercept_ yHat = ourModel.predict(X) print(f"R² Score: {rSq}") print(f"Coefficients: {slope}") print(f"Intercept: {intercept}") # 可视化(因多特征无法直接画二维散点,示例展示能量特征与偏好的实际/预测关系) plt.scatter(energy_flat, y, alpha=0.3, label='Actual') plt.scatter(energy_flat, yHat, alpha=0.3, color='orange', label='Predicted') plt.xlabel('Energy Rating') plt.ylabel('Preference Rating') plt.title('Predicting Art Preference from Energy and Demographic') plt.legend() plt.show()
关键说明
np.repeat(demo, repeats=91, axis=0):确保每个用户的6项人口统计数据被重复91次,对应该用户对91件作品的所有评分样本。np.column_stack:将一维的能量评分和二维的人口统计数据拼接成完整的特征矩阵,符合sklearn对输入特征的要求(每行一个样本,每列一个特征)。
内容的提问来源于stack exchange,提问作者Angie Liu
相关产品推荐
相关产品推荐

