使用Sklearn拟合线性回归模型时遇ValueError报错求助
问题:Sklearn LinearRegression拟合Pandas数据时出现维度错误
原代码
features =["floors", "waterfront", "lat", "bedrooms", "sqft_basement", "view", "bathrooms", "sqft_living15", "sqft_above", "grade", "sqft_living"] lm3 = LinearRegression() lm3.fit(features,["prices"]) lm3.score(features,B)
报错信息
ValueError: Expected 2D array, got 1D array instead: array=['floors' 'waterfront' 'lat' 'bedrooms' 'sqft_basement' 'view' 'bathrooms' 'sqft_living15' 'sqft_above' 'grade' 'sqft_living']. Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
完整报错栈:
ValueError Traceback (most recent call last) Cell In[27], line 3 1 features =["floors", "waterfront","lat" ,"bedrooms" ,"sqft_basement" ,"view" ,"bathrooms","sqft_living15","sqft_above","grade","sqft_living"] 2 lm3 = LinearRegression() ----> 3 lm3.fit(features,["prices"]) 4 lm3.score(features,B) File D:\python\Lib\site-packages\sklearn\linear_model\_base.py:648, in LinearRegression.fit(self, X, y, sample_weight) 644 n_jobs_ = self.n_jobs 646 accept_sparse = False if self.positive else ["csr", "csc", "coo"] ---> 648 X, y = self._validate_data( 649 X, y, accept_sparse=accept_sparse, y_numeric=True, multi_output=True 650 ) 652 sample_weight = _check_sample_weight( 653 sample_weight, X, dtype=X.dtype, only_non_negative=True 654 ) 656 X, y, X_offset, y_offset, X_scale = _preprocess_data( 657 X, 658 y, (...) 661 sample_weight=sample_weight, 662 ) File D:\python\Lib\site-packages\sklearn\base.py:584, in BaseEstimator._validate_data(self, X, y, reset, validate_separately, **check_params) 582 y = check_array(y, input_name="y", **check_y_params) 583 else: ---> 584 X, y = check_X_y(X, y, **check_params) 585 out = X, y 587 if not no_val_X and check_params.get("ensure_2d", True): File D:\python\Lib\site-packages\sklearn\utils\validation.py:1106, in check_X_y(X, y, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, multi_output, ensure_min_samples, ensure_min_features, y_numeric, estimator) 1101 estimator_name = _check_estimator_name(estimator) 1102 raise ValueError( 1103 f"{estimator_name} requires y to be passed, but the target y is None" 1104 ) -> 1106 X = check_array( 1107 X, 1108 accept_sparse=accept_sparse, 1109 accept_large_sparse=accept_large_sparse, 1110 dtype=dtype, 1111 order=order, 1112 copy=copy, 1113 force_all_finite=force_all_finite, 1114 ensure_2d=ensure_2d, 1115 allow_nd=allow_nd, 1116 ensure_min_samples=ensure_min_samples, 1117 ensure_min_features=ensure_min_features, 1118 estimator=estimator, 1119 input_name="X", 1120 ) 1122 y = _check_y(y, multi_output=multi_output, y_numeric=y_numeric, estimator=estimator) 1124 check_consistent_length(X, y) File D:\python\Lib\site-packages\sklearn\utils\validation.py:902, in check_array(array, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, ensure_min_samples, ensure_min_features, estimator, input_name) 900 # If input is 1D raise error 901 if array.ndim == 1: ---> 902 raise ValueError( 903 "Expected 2D array, got 1D array instead:\narray={}.\n" 904 "Reshape your data either using array.reshape(-1, 1) if " 905 "your data has a single feature or array.reshape(1, -1) " 906 "if it contains a single sample.".format(array) 907 ) 909 if dtype_numeric and array.dtype.kind in "USV": 910 raise ValueError( 911 "dtype='numeric' is not compatible with arrays of bytes/strings." 912 "Convert your data to numeric values explicitly instead." 913 ) ValueError: Expected 2D array, got 1D array instead: array=['floors' 'waterfront' 'lat' 'bedrooms' 'sqft_basement' 'view' 'bathrooms' 'sqft_living15' 'sqft_above' 'grade' 'sqft_living']. Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
问题原因
直接将特征名称列表传给了fit()和score()方法,而非从Pandas DataFrame中提取对应特征的数值数据。Sklearn的LinearRegression.fit()要求:
- 第一个参数
X是二维特征矩阵(每行一个样本,每列一个特征) - 第二个参数
y是一维目标变量数组(每个样本对应的标签)
原代码中,features是字符串列表,属于一维数组,不符合输入维度要求;["prices"]也是列表,不是目标变量的数值序列。
解决方案
假设数据存储在名为df的Pandas DataFrame中,需从df提取特征列和目标列后传入模型:
from sklearn.linear_model import LinearRegression import pandas as pd # 假设数据存储在df中 features = ["floors", "waterfront", "lat", "bedrooms", "sqft_basement", "view", "bathrooms", "sqft_living15", "sqft_above", "grade", "sqft_living"] lm3 = LinearRegression() # 提取特征矩阵X和目标变量y X = df[features] y = df["prices"] # 拟合模型 lm3.fit(X, y) # 计算评分(替换原代码中未定义的B为真实目标变量y) model_score = lm3.score(X, y) print(model_score)
关键说明
df[features]返回DataFrame,是符合要求的二维特征矩阵df["prices"]返回Series,是一维目标变量数组,满足Sklearn输入要求- 原代码中
B未定义,需替换为真实目标变量数据
内容的提问来源于stack exchange,提问作者Legofan35664
相关产品推荐
相关产品推荐

