You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何正确使用dropna处理NaN值以解决LinearRegression建模报错

解决LinearRegression输入含NaN的问题

你的报错出现在dummy编码分支的model_d.fit(X_train_d,y_train_d),说明df2生成的X_d中存在NaN值,且你之前的处理存在两个核心问题:

  1. 仅处理了第一个模型的X,未同步处理对应的y,会导致特征和标签的索引不匹配;
  2. 完全没处理dummy分支的df2数据,这才是报错的直接源头。

下面提供两种可行的解决方案:

方案一:删除含NaN的样本(dropna正确用法)

需要同步清理特征(X)和标签(y),避免出现行数不一致的问题,同时对两个数据集(df和df2)都做处理:

from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
import pandas as pd

# Integer encoding 分支数据清理
X = df.drop(labels=['Category','Rating','Genres','Genres_c'], axis=1)
y = df.Rating

# 合并特征和标签后统一删除含NaN的行,保证索引一致
combined_df = pd.concat([X, y], axis=1).dropna(axis=0, how='any')
X_clean = combined_df.drop('Rating', axis=1)
y_clean = combined_df['Rating']

X_train, X_test, y_train, y_test = train_test_split(X_clean, y_clean, test_size=0.30)

model = LinearRegression()
model.fit(X_train, y_train)
Results = model.predict(X_test)

# 结果表格生成(原代码不变)
resultsdf = pd.DataFrame()
resultsdf = resultsdf.from_dict(Evaluationmatrix_dict(y_test, Results), orient='index')
resultsdf = resultsdf.transpose()

# Dummy encoding 分支数据清理
X_d = df2.drop(labels=['Rating','Genres','Category_c','Genres_c'], axis=1)
y_d = df2.Rating

# 同样同步清理df2的特征和标签
combined_df2 = pd.concat([X_d, y_d], axis=1).dropna(axis=0, how='any')
X_d_clean = combined_df2.drop('Rating', axis=1)
y_d_clean = combined_df2['Rating']

X_train_d, X_test_d, y_train_d, y_test_d = train_test_split(X_d_clean, y_d_clean, test_size=0.30)

model_d = LinearRegression()
model_d.fit(X_train_d, y_train_d)  # 此时X_train_d无NaN,可正常训练
Results_d = model_d.predict(X_test_d)

resultsdf = resultsdf.append(Evaluationmatrix_dict(y_test_d, Results_d, name='Linear - Dummy'), ignore_index=True)

方案二:填充NaN值(保留所有样本)

如果不想删除样本,可以用SimpleImputer填充缺失值,结合Pipeline避免数据泄露:

from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.impute import SimpleImputer
from sklearn.pipeline import make_pipeline
import pandas as pd

# Integer encoding 分支用管道处理
X = df.drop(labels=['Category','Rating','Genres','Genres_c'], axis=1)
y = df.Rating
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.30)

# 管道:先用均值填充NaN,再训练线性回归
model = make_pipeline(SimpleImputer(strategy='mean'), LinearRegression())
model.fit(X_train, y_train)
Results = model.predict(X_test)

# 结果表格生成(原代码不变)
resultsdf = pd.DataFrame()
resultsdf = resultsdf.from_dict(Evaluationmatrix_dict(y_test, Results), orient='index')
resultsdf = resultsdf.transpose()

# Dummy encoding 分支同理
X_d = df2.drop(labels=['Rating','Genres','Category_c','Genres_c'], axis=1)
y_d = df2.Rating
X_train_d, X_test_d, y_train_d, y_test_d = train_test_split(X_d, y_d, test_size=0.30)

model_d = make_pipeline(SimpleImputer(strategy='mean'), LinearRegression())
model_d.fit(X_train_d, y_train_d)
Results_d = model_d.predict(X_test_d)

resultsdf = resultsdf.append(Evaluationmatrix_dict(y_test_d, Results_d, name='Linear - Dummy'), ignore_index=True)

说明:SimpleImputer的strategy参数可根据需求调整,比如median(中位数)、most_frequent(众数)等。

内容的提问来源于stack exchange,提问作者Rupesh Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 15:13:08