Scikit-Learn中用Datetime做线性回归预测的错误排查与实现方法
问题:如何以datetime为单一特征用线性回归实现能耗预测?
我有包含row_id、datetime、energy字段的数据集,其中datetime为object类型,energy为float64类型,希望以datetime作为唯一特征预测能耗值。
我最初编写的代码:
train['datetime'] = pd.to_datetime(train['datetime']) X = train.iloc[:,0] y = train.iloc[:,-1]
运行后出现错误:
ValueError: Expected 2D array, got 1D array instead: array=['2008-03-01T00:00:00.000000000' '2008-03-01T01:00:00.000000000' '2008-03-01T02:00:00.000000000' ... '2018-12-31T21:00:00.000000000' '2018-12-31T22:00:00.000000000' '2018-12-31T23:00:00.000000000']. Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
于是我调整了数据形状:
X = np.array(X).reshape(-1,1) y = np.array(y).reshape(-1,1) from sklearn.linear_model import LinearRegression model_1 = LinearRegression() model_1.fit(X,y) test = pd.to_datetime(test['datetime']) test = np.array(test).reshape(-1,1) predictions = model_1.predict(test)
模型拟合时未报错,但调用predict方法传入测试数据时,出现类型错误:
TypeError: The DType <class 'numpy.dtype[datetime64]'> could not be promoted by <class 'numpy.dtype[float64]'>. This means that no common DType exists for the given inputs. For example they cannot be stored in a single array unless the dtype is `object`. The full list of DTypes is: (<class 'numpy.dtype[datetime64]'>, <class 'numpy.dtype[float64]'>)
解决方案
错误原因
sklearn的线性回归模型无法直接处理datetime64类型的特征,它要求输入必须是数值型数据。你之前拟合时未报错是因为模型内部可能做了隐式类型转换,但预测时测试数据的datetime类型与模型期望的数值类型不匹配,导致类型冲突。
解决步骤
1. 将datetime转换为数值型特征
把datetime转换成时间戳(单位秒/毫秒),或者相对于起始时间的天数/小时数,将日期类型转为float64数值,符合模型输入要求。
示例代码:
import pandas as pd import numpy as np from sklearn.linear_model import LinearRegression # 处理训练数据 train['datetime'] = pd.to_datetime(train['datetime']) # 将datetime转换为时间戳(秒),并转为2D数组 X_train = train['datetime'].apply(lambda x: x.timestamp()).values.reshape(-1, 1) y_train = train['energy'].values.reshape(-1, 1) # 训练线性回归模型 model_1 = LinearRegression() model_1.fit(X_train, y_train) # 处理测试数据,必须和训练数据用相同的转换方式 test['datetime'] = pd.to_datetime(test['datetime']) X_test = test['datetime'].apply(lambda x: x.timestamp()).values.reshape(-1, 1) # 生成预测结果 predictions = model_1.predict(X_test)
2. 时间序列预测的补充说明
如果你的场景是时间序列预测,单纯用时间戳的线性回归只能捕捉整体趋势(比如能耗随时间递增/递减),无法处理周期性波动(比如每日、每周的能耗规律)。如果需要考虑周期性,可以额外提取时间特征:
- 提取小时、星期几、月份等特征加入模型
- 示例代码:
# 提取时间特征 train['timestamp'] = train['datetime'].apply(lambda x: x.timestamp()) train['hour'] = train['datetime'].dt.hour train['dayofweek'] = train['datetime'].dt.dayofweek train['month'] = train['datetime'].dt.month # 使用多特征训练模型 X_train = train[['timestamp', 'hour', 'dayofweek', 'month']].values model_1.fit(X_train, y_train) # 测试数据同样提取这些特征 test['timestamp'] = test['datetime'].apply(lambda x: x.timestamp()) test['hour'] = test['datetime'].dt.hour test['dayofweek'] = test['datetime'].dt.dayofweek test['month'] = test['datetime'].dt.month X_test = test[['timestamp', 'hour', 'dayofweek', 'month']].values predictions = model_1.predict(X_test)
内容的提问来源于stack exchange,提问作者Rama
相关产品推荐
相关产品推荐

