Andrew Ng机器学习课程可视化报错:x和y尺寸不匹配求助
散点图绘制报错:x与y维度不匹配问题解决
问题背景
我正在学习Andrew Ng的机器学习课程,尝试在Python Spyder中复现课程中的数据可视化。
我的数据数组如下:
[['Size' 'Bedrooms' 'Floors' 'Age Home' 'Price'] ['952' '2' '1' '65' '271.5'] ['1244' '3' '2' '64' '232'] ['1974' '3' '2' '17' '509.8']]
想要将Size、Bedrooms、Floors、Age Home分别作为x轴,Price作为y轴绘制散点图,编写的代码如下:
arr = np.loadtxt(r"", delimiter=",", dtype=str) x_train = arr[1:, 2] y_train = arr[1:, 4] x_features = ["Size","Bedrooms","Floors","Age Home"] print(x_train) print(y_train) fig,ax=plt.subplots(1, 4, figsize=(12, 3), sharey=True) for i in range(len(ax)): ax[i].scatter(x_train[1: i],y_train) ax[i].set_xlabel(x_features[i]) ax[0].set_ylabel("Price (1000's)") plt.show()
但持续报错:
raise ValueError("x and y must be the same size") ValueError: x and y must be the same size
问题分析与修复
报错核心是x和y的长度不匹配,代码存在两个关键问题:
- 仅提取了单一列作为x_train,没有覆盖需要的4个特征列
- 切片
x_train[1: i]语法错误,循环时会生成长度与y_train不一致的数组
修改后的完整代码
import numpy as np import matplotlib.pyplot as plt # 替换为你的数据文件路径 arr = np.loadtxt(r"你的文件路径", delimiter=",", dtype=str) # 提取所有4个特征列和价格列,并转换为数值类型 x_features = ["Size","Bedrooms","Floors","Age Home"] x_data = arr[1:, 0:4].astype(float) y_data = arr[1:, 4].astype(float) fig, ax = plt.subplots(1, 4, figsize=(12, 3), sharey=True) for i in range(len(ax)): # 取对应特征列作为x轴,确保与y轴长度一致 ax[i].scatter(x_data[:, i], y_data) ax[i].set_xlabel(x_features[i]) ax[0].set_ylabel("Price (1000's)") plt.tight_layout() # 优化子图布局,避免标签重叠 plt.show()
关键修改说明
- 修正x轴数据源:提取数组中索引0到3的4个特征列,而非单一列,同时将字符串转换为浮点型,保证绘图数据的正确性
- 修正循环取值逻辑:用
x_data[:, i]每次取对应特征的所有数据,确保和y_data长度完全匹配 - 补全库导入:明确导入numpy和matplotlib.pyplot,避免环境依赖报错
- 优化布局:添加
plt.tight_layout()自动调整子图间距,提升可视化效果
内容的提问来源于stack exchange,提问作者spacemen
相关产品推荐
相关产品推荐

