解决TypeError: 'numpy.ndarray' object is not callable错误:房价预测线性回归问题
房价预测线性回归代码报错:TypeError: 'numpy.ndarray' object is not callable
问题背景
正在开展房价预测的简单线性回归任务,采用RMSE作为评估指标,编写代码后触发TypeError: 'numpy.ndarray' object is not callable错误,无法定位问题来源。
原代码
# A function that take one input of the dataset and return the RMSE (of the test data), and the intercept and coefficient def simple_linear_model(train, test, input_feature): regr = linear_model.LinearRegression() # Create a linear regression object regr.fit(train.values(columns = [input_feature]), train.as_matrix(columns = ['Price'])) # Train the model RMSE = mean_squared_error(test.values(columns = ['Price']), regr.predict(test.values(columns = [input_feature])))**0.5 # Calculate the RMSE on test data return RMSE, regr.intercept_[0], regr.coef_[0][0] input_list = data.columns.values.tolist() # list of column name input_list.remove('Price') simple_linear_result = pd.DataFrame(columns = ['feature', 'RMSE', 'intercept', 'coefficient']) # loop that calculate the RMSE of the test data for each input for p in input_list: RMSE, w1, w0 = simple_linear_model(train_data, test_data, p) simple_linear_result = simple_linear_result.append({'feature':p, 'RMSE':RMSE, 'intercept':w0, 'coefficient': w1} ,ignore_index=True) simple_linear_result.sort_values('RMSE').head(10) # display the 10 best estimators
报错信息
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) Input In [12], in <cell line: 6>() 5 # loop that calculate the RMSE of the test data for each input 6 for p in input_list: ----> 7 RMSE, w1, w0 = simple_linear_model(train_data, test_data, p) 8 simple_linear_result = simple_linear_result.append({'feature':p, 'RMSE':RMSE, 'intercept':w0, 'coefficient': w1} 9 ,ignore_index=True) 10 simple_linear_result.sort_values('RMSE').head(10) Input In [11], in simple_linear_model(train, test, input_feature) 2 def simple_linear_model(train, test, input_feature): 3 regr = linear_model.LinearRegression() # Create a linear regression object ----> 4 regr.fit(train.values(columns = [input_feature]), train.as_matrix(columns = ['Price'])) # Train the model 5 RMSE = mean_squared_error(test.values(columns = ['Price']), 6 regr.predict(test.values(columns = [input_feature])))**0.5 # Calculate the RMSE on test data 7 return RMSE, regr.intercept_[0], regr.coef_[0][0] TypeError: 'numpy.ndarray' object is not callable
错误原因
错误根源在以下两处:
train.values(columns = [input_feature]):values是pandas DataFrame的属性,不是方法,直接访问train.values会返回一个numpy数组。你试图用调用方法的方式(加括号传参数)去使用它,就会触发错误——numpy数组不能被当作函数调用。train.as_matrix(columns = ['Price']):as_matrix是已被废弃的pandas方法,且同样错误地用了类似方法调用的传参方式,另外即使正确使用,它返回的也是numpy数组,后续操作也会有问题。
此外,线性回归模型的输入特征需要是二维数组,直接取单列的values会得到一维数组,也会导致模型拟合报错。
修正后的代码
from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error import pandas as pd # 定义线性回归模型函数,返回测试集RMSE、截距和系数 def simple_linear_model(train, test, input_feature): regr = LinearRegression() # 提取特征:转为二维数组 X_train = train[[input_feature]].values # 提取目标变量:转为二维数组 y_train = train[['Price']].values regr.fit(X_train, y_train) # 计算测试集RMSE X_test = test[[input_feature]].values y_test = test[['Price']].values y_pred = regr.predict(X_test) RMSE = mean_squared_error(y_test, y_pred)**0.5 return RMSE, regr.intercept_[0], regr.coef_[0][0] input_list = data.columns.values.tolist() input_list.remove('Price') simple_linear_result = pd.DataFrame(columns=['feature', 'RMSE', 'intercept', 'coefficient']) # 遍历所有特征计算结果 for p in input_list: RMSE, w1, w0 = simple_linear_model(train_data, test_data, p) simple_linear_result = simple_linear_result.append( {'feature':p, 'RMSE':RMSE, 'intercept':w0, 'coefficient': w1}, ignore_index=True ) # 展示RMSE最小的10个特征模型 simple_linear_result.sort_values('RMSE').head(10)
关键修正点
- 用
train[[input_feature]].values替代错误的train.values(columns=...),确保特征是二维数组 - 用
train[['Price']].values替代train.as_matrix(columns=...),同时保证目标变量是二维数组 - 拆分预测和RMSE计算步骤,让代码更清晰易维护
内容的提问来源于stack exchange,提问作者user20273524
相关产品推荐
相关产品推荐

