LinearRegression拟合报错:期望2D数组却得到1D数组的解决方法咨询
问题:LinearRegression拟合时提示"Expected 2D array, got 1D array instead"
我用LinearRegression模型预测年度'Value'数值,从CSV导入的数据如下(df.head()结果):
LOCATION INDICATOR SUBJECT MEASURE FREQUENCY TIME Value Flag Codes 0 IRL LTINT TOT PC_PA M 1.167610e+18 4.04 NaN 1 IRL LTINT TOT PC_PA M 1.170288e+18 4.07 NaN 2 IRL LTINT TOT PC_PA M 1.172707e+18 3.97 NaN 3 IRL LTINT TOT PC_PA M 1.175386e+18 4.19 NaN 4 IRL LTINT TOT PC_PA M 1.177978e+18 4.32 NaN
原TIME列格式为YYYY-DD-MM,我执行了以下代码处理该列:
df['TIME'] = pd.to_datetime(df['TIME']) df['TIME'] = pd.to_numeric(df['TIME']) df['TIME'] = df['TIME'].astype(float)
随后编写拟合代码:
X=df['TIME'] Y=df['Value'] X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.4, random_state=101) lm= LinearRegression()
执行lm.fit(X_train, Y_train)时出现报错:
ValueError Traceback (most recent call last) Cell In[12], line 1 ----> 1 lm.fit(X_train, Y_train) File ~\anaconda3\Lib\site-packages\sklearn\base.py:1151, in _fit_context.<locals>.decorator.<locals>.wrapper(estimator, *args, **kwargs) 1144 estimator._validate_params() 1146 with config_context( 1147 skip_parameter_validation=( 1148 prefer_skip_nested_validation or global_skip_validation 1149 ) 1150 ): -> 1151 return fit_method(estimator, *args, **kwargs) File ~\anaconda3\Lib\site-packages\sklearn\linear_model\_base.py:678, in LinearRegression.fit(self, X, y, sample_weight) 674 n_jobs_ = self.n_jobs 676 accept_sparse = False if self.positive else ["csr", "csc", "coo"] -> 678 X, y = self._validate_data( 679 X, y, accept_sparse=accept_sparse, y_numeric=True, multi_output=True 680 ) 682 has_sw = sample_weight is not None 683 if has_sw: File ~\anaconda3\Lib\site-packages\sklearn\base.py:621, in BaseEstimator._validate_data(self, X, y, reset, validate_separately, cast_to_ndarray, **check_params) 619 y = check_array(y, input_name="y", **check_y_params) 620 else: -> 621 X, y = check_X_y(X, y, **check_params) 622 out = X, y 624 if not no_val_X and check_params.get("ensure_2d", True): File ~\anaconda3\Lib\site-packages\sklearn\utils\validation.py:1147, in check_X_y(X, y, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, multi_output, ensure_min_samples, ensure_min_features, y_numeric, estimator) 1142 estimator_name = _check_estimator_name(estimator) 1143 raise ValueError( 1144 f"{estimator_name} requires y to be passed, but the target y is None" 1145 ) -> 1147 X = check_array( 1148 X, 1149 accept_sparse=accept_sparse, 1150 accept_large_sparse=accept_large_sparse, 1151 dtype=dtype, 1152 order=order, 1153 copy=copy, 1154 force_all_finite=force_all_finite, 1155 ensure_2d=ensure_2d, 1156 allow_nd=allow_nd, 1157 ensure_min_samples=ensure_min_samples, 1158 ensure_min_features=ensure_min_features, 1159 estimator=estimator, 1160 input_name="X", 1161 ) 1163 y = _check_y(y, multi_output=multi_output, y_numeric=y_numeric, estimator=estimator) 1165 check_consistent_length(X, y) File ~\anaconda3\Lib\site-packages\sklearn\utils\validation.py:940, in check_array(array, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, ensure_min_samples, ensure_min_features, estimator, input_name) 938 # If input is 1D raise error 939 if array.ndim == 1: -> 940 raise ValueError( 941 "Expected 2D array, got 1D array instead:\narray={}.\n" 942 "Reshape your data either using array.reshape(-1, 1) if " 943 "your data has a single feature or array.reshape(1, -1) " 944 "if it contains a single sample.".format(array) 945 ) 947 if dtype_numeric and hasattr(array.dtype, "kind") and array.dtype.kind in "USV": 948 raise ValueError( 949 "dtype='numeric' is not compatible with arrays of bytes/strings." 950 "Convert your data to numeric values explicitly instead." 951 ) ValueError: Expected 2D array, got 1D array instead: array=[1.3437792e+18 1.6198272e+18 1.3596768e+18 ... 1.3596768e+18 1.5805152e+18 1.5751584e+18]. Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
解决方案
错误核心原因:scikit-learn的LinearRegression要求输入的特征矩阵X必须是二维数组(形状为(n_samples, n_features)),但你传入的X_train是一维的Series(或一维numpy数组)。
以下三种方法可解决问题:
方法1:用双括号选取特征列,直接生成二维DataFrame
修改X的定义,使用df[['TIME']]代替df['TIME'],双括号会返回二维的DataFrame,符合模型输入要求:
X = df[['TIME']] # 返回形状为(n,1)的DataFrame Y = df['Value'] X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.4, random_state=101) lm = LinearRegression() lm.fit(X_train, Y_train) # 可正常执行
方法2:用reshape将一维数组转为二维
在拟合时或提前对X进行reshape处理:
X = df['TIME'] Y = df['Value'] X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.4, random_state=101) lm = LinearRegression() # 拟合时对X_train进行reshape lm.fit(X_train.values.reshape(-1, 1), Y_train)
或提前处理X:
X = df['TIME'].values.reshape(-1, 1) # 直接转为二维数组 Y = df['Value'] X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.4, random_state=101) lm = LinearRegression() lm.fit(X_train, Y_train)
方法3:用sklearn转换器确保维度(适合复杂流程)
如果后续有更多特征处理步骤,可使用FunctionTransformer统一处理维度:
from sklearn.preprocessing import FunctionTransformer reshape_transformer = FunctionTransformer(lambda x: x.reshape(-1, 1)) X = reshape_transformer.fit_transform(df['TIME']) Y = df['Value'] X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.4, random_state=101) lm = LinearRegression() lm.fit(X_train, Y_train)
内容的提问来源于stack exchange,提问作者Najoua
相关产品推荐
相关产品推荐

