You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多变量回归报错ValueError:样本数不一致问题排查求助

多变量回归样本数不匹配问题排查

问题背景

数据集来自300人对91件艺术品的调查:

  • 前91列:1-7分的偏好评分
  • 第92-182列:作品能量评分
  • 第216-221列:人口统计数据

目标是基于能量评分和人口统计数据构建回归模型预测偏好评分,单变量回归可正常运行,但多变量回归执行时报错。

原代码

rate = data[:,0:91] #isolate preference ratings
energy = data[:,91:182] #isolate energy ratings
demo = data[:,215:221] #isolate demographic ratings
demo[np.isnan(demo)] = 4.68
comb = np.concatenate((energy,demo),axis=1)
x = comb.reshape(-1,1)
y = rate.reshape(-1,1)
ourModel = LinearRegression().fit(x, y)
rSq = ourModel.score(x,y)
slope = ourModel.coef_
intercept = ourModel.intercept_
yHat = slope * x + intercept
print(rSq)

plt.plot(x,y,'o',)
plt.xlabel('Energy + Demographic') 
plt.ylabel('Preference')  
plt.plot(x,yHat,color='orange',linewidth=3)
plt.title('Predicting Art Preference from Energy and Demographic') 
plt.show()

报错信息

ValueError: Found input variables with inconsistent numbers of samples: [29100, 27300]

报错原因

样本数不匹配的根源是数据重塑逻辑错误:

  • comb是将300行的energy(91列)和demo(6列)按列拼接,得到(300, 97)的数组,reshape(-1,1)后变成(300*97, 1) = (29100,1),这是错误的样本数。
  • rate是(300,91)的数组,reshape(-1,1)后得到(300*91,1)=(27300,1),这是正确的样本数(每个用户对每件作品的评分对应一个样本)。

核心逻辑错误:没有建立每个偏好评分对应该作品能量评分+用户人口统计数据的正确映射,而是错误地把所有特征列强行压成单特征样本,导致样本数不匹配。

修正方案

需要将数据整理为:每个样本对应「1个偏好评分」+「该作品的能量评分+对应用户的人口统计数据」,具体步骤:

  1. 展开rate为一维数组(27300个样本),保持每个评分的独立性。
  2. 展开energy为一维数组(每个作品的能量评分对应到用户的评分样本)。
  3. 将demo数据重复91次(每个用户的人口统计数据要对应他对91件作品的所有评分),得到和样本数匹配的(27300,6)数组。
  4. 拼接展开后的energy和重复后的demo,得到特征矩阵X,形状为(27300,7),与y的样本数完全匹配。

修正后代码

import numpy as np
from sklearn.linear_model import LinearRegression
import matplotlib.pyplot as plt

# 提取数据
rate = data[:,0:91] # 偏好评分 (300,91)
energy = data[:,91:182] # 能量评分 (300,91)
demo = data[:,215:221] # 人口统计数据 (300,6)

# 填充缺失值
demo[np.isnan(demo)] = 4.68

# 处理特征和标签,建立正确的样本映射
y = rate.flatten() # 展开为(27300,)的一维数组,每个元素是一个偏好评分样本
energy_flat = energy.flatten() # 展开能量评分为(27300,)
# 将人口统计数据重复91次,每个用户的demo对应91个作品评分
demo_repeated = np.repeat(demo, repeats=91, axis=0) # 形状变为(27300,6)

# 拼接特征:能量评分(1列) + 人口统计(6列)
X = np.column_stack((energy_flat, demo_repeated)) # 形状(27300,7)

# 训练模型
ourModel = LinearRegression().fit(X, y)
rSq = ourModel.score(X, y)
slope = ourModel.coef_
intercept = ourModel.intercept_
yHat = ourModel.predict(X)

print(f"R² Score: {rSq}")
print(f"Coefficients: {slope}")
print(f"Intercept: {intercept}")

# 可视化(因多特征无法直接画二维散点,示例展示能量特征与偏好的实际/预测关系)
plt.scatter(energy_flat, y, alpha=0.3, label='Actual')
plt.scatter(energy_flat, yHat, alpha=0.3, color='orange', label='Predicted')
plt.xlabel('Energy Rating') 
plt.ylabel('Preference Rating')  
plt.title('Predicting Art Preference from Energy and Demographic')
plt.legend()
plt.show()

关键说明

  • np.repeat(demo, repeats=91, axis=0):确保每个用户的6项人口统计数据被重复91次,对应该用户对91件作品的所有评分样本。
  • np.column_stack:将一维的能量评分和二维的人口统计数据拼接成完整的特征矩阵,符合sklearn对输入特征的要求(每行一个样本,每列一个特征)。

内容的提问来源于stack exchange,提问作者Angie Liu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 17:10:16