线性回归调用StandardScaler时报错float()参数不能为NaTType求助
问题原因与解决方案
报错触发在自变量的缩放环节,NaTType是Pandas中 datetime 类型的空值标识,说明自变量数据集SourceData_train_independent里存在包含空值的时间类型列,StandardScaler只能处理数值类型输入,无法解析NaT值,仅修改数据类型不能直接解决问题,需要先完成异常列的排查和特征处理。
解决步骤
- 第一步:定位存在NaT的列
运行如下代码排查自变量中的时间列空值情况:
# 查看所有时间类型列的空值数量 print(SourceData_train_independent.select_dtypes(include=['datetime64', 'datetime64[ns]']).isna().sum())
- 第二步:根据业务需求处理异常列
- 若该时间列不需要纳入模型:直接删除对应列即可,测试集自变量也要同步删除相同列
# 替换为你查到的异常时间列名 drop_cols = ['异常列名1', '异常列名2'] SourceData_train_independent = SourceData_train_independent.drop(drop_cols, axis=1) SourceData_test_independent = SourceData_test_independent.drop(drop_cols, axis=1) - 若需要保留时间特征:将时间列转换为数值型特征后再使用,常见转换方式包括拆分出年、月、日、星期几等离散特征,或转换为距离基准日期的天数,空值可以用列的中位数、众数填充,也可以单独标记为特殊数值:
import pandas as pd # 填充空时间 SourceData_train_independent['时间列名'] = SourceData_train_independent['时间列名'].fillna(SourceData_train_independent['时间列名'].median()) # 转换为距离基准日期的天数 SourceData_train_independent['time_diff'] = (SourceData_train_independent['时间列名'] - pd.to_datetime('2000-01-01')).dt.days # 删除原时间列 SourceData_train_independent = SourceData_train_independent.drop('时间列名', axis=1) # 测试集做完全相同的处理
- 若该时间列不需要纳入模型:直接删除对应列即可,测试集自变量也要同步删除相同列
额外注意事项
当前代码存在严重的数据错误,给测试集Sale Price赋值时,错误调用了训练集的Sale Price字段,会导致后续模型评估结果完全失真,需要修正为:
Test_Data["Sale Price"] = Test_Data["Sale Price"].astype(str).str.strip().replace("",0).astype(float)
处理完以上步骤后再运行StandardScaler的相关代码即可正常执行。
内容的提问来源于stack exchange,提问作者The_Bandit
相关产品推荐
相关产品推荐

