已完成Train-Test Split仍报NameError: name 'X_train' is not defined求助
NameError: 'X_train'未定义问题排查与解决
问题背景
已成功完成训练测试集划分,输出的数据集维度验证了划分有效,但运行建模代码时触发NameError,提示X_train未定义,重新执行划分代码后问题依旧。
相关代码与报错信息
训练测试集划分代码及输出
# Test-train split from sklearn.model_selection import train_test_split X = df.iloc[:,0:5] y = df['charges'] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=1) print(X_train.shape) #(3500, 5) print(X_test.shape) #(1500, 5) print(y_train.shape) #(3500, ) print(y_test.shape) #(1500, )
输出结果:
(936, 5) (402, 5) (936,) (402,)
建模代码及报错
# Modeling from sklearn.linear_model import LinearRegression from sklearn import metrics from sklearn.metrics import mean_squared_error reg = LinearRegression() reg.fit(X_train,y_train) print("The intercept term of the linear model:", reg.intercept_) print("The coefficients of the linear model:", reg.coef_)
报错信息:
--------------------------------------------------------------------------- NameError Traceback (most recent call last) <ipython-input-15-c8fa5d4c8df7> in <cell line: 8>() ----> reg.fit(X_train,y_train) print("The intercept term of the linear model:", reg.intercept_) NameError: name 'X_train' is not defined
排查与解决方向
- 确认代码执行上下文一致性:如果使用Jupyter Notebook,确保划分代码和建模代码在同一个内核会话中执行,且划分代码的单元格已成功运行(无报错)。避免在不同内核或重启内核后只运行建模代码。
- 检查变量作用域与拼写:确认建模代码中
X_train的拼写与划分代码完全一致,无大小写错误或字符遗漏;同时检查划分代码后是否有其他代码删除或覆盖了X_train(比如del X_train操作)。 - 重启内核并顺序执行:Jupyter环境偶尔会出现变量缓存异常,重启内核后按顺序执行所有代码(先运行数据集加载、划分代码,再运行建模代码),排除变量污染问题。
- 验证变量存在性:在建模代码前添加
print('X_train' in locals())或print(X_train),确认X_train变量是否真的存在于当前环境中。
内容的提问来源于stack exchange,提问作者Alex Tsai
相关产品推荐
相关产品推荐

