使用statsmodels后向消元时触发TypeError: ufunc 'isfinite'输入类型不支持
Hey there, let's break down what's going wrong here and get your backward elimination working smoothly:
First off, the biggest mistake in your code is that you've mixed up the arguments for smf.OLS() — you passed your feature matrix readX to endog (which is supposed to be your target variable, readY), and your optimized feature set to exog. endog stands for "endogenous variable" (the one you're predicting), while exog is the "exogenous" feature matrix. That critical mix-up is driving most of the error.
On top of that, there's a data type issue: the output from ColumnTransformer might be an object-type array, and when you appended the integer-type ones array, it created a mixed-type matrix that np.isfinite (used internally by statsmodels) can't handle.
Step-by-Step Fixes:
- Correct the OLS Argument Order: Swap the
endogandexogparameters soendog=readY(your target profit variable) andexog=optimisedX(your feature matrix with the constant term). - Standardize Data Types: Convert your transformed feature matrix to
float64explicitly, and use a float-type ones array when appending the constant term — this avoids type mismatches. - Clean Up the Feature Matrix: Ensure the output from
ColumnTransformeris converted to a proper numeric array right after transformation.
Fixed Full Code:
import pandas as pd import matplotlib.pyplot as plt import numpy as np from sklearn.linear_model import LinearRegression from sklearn.model_selection import train_test_split from sklearn.preprocessing import OneHotEncoder from sklearn.compose import ColumnTransformer import statsmodels.api as smf data = pd.read_csv('F:/Py Projects/ML_Dataset/50_Startups.csv') dataSlice = data.head(10) # Extract features and target variable readX = data.iloc[:, :4].values readY = data.iloc[:, 4].values # Encode categorical feature (State) transformer = ColumnTransformer( transformers=[("OneHot", OneHotEncoder(), [3])], remainder='passthrough' ) # Convert to float64 immediately to avoid type issues readX = transformer.fit_transform(readX).astype(np.float64) readX = readX[:, 1:] # Avoid dummy variable trap # Optional train-test split (backward elimination often uses full dataset) trainX, testX, trainY, testY = train_test_split(readX, readY, test_size=0.2, random_state=0) lreg = LinearRegression() lreg.fit(trainX, trainY) predY = lreg.predict(testX) # Append constant term with matching float type readX = np.append(arr=np.ones((50, 1), dtype=np.float64), values=readX, axis=1) optimisedX = readX[:, [0, 1, 2, 3, 4, 5]] # Fixed OLS call: endog is target, exog is features ols = smf.OLS(endog=readY, exog=optimisedX).fit() print(ols.summary())
Why This Works:
- Argument Fix: By passing
readYtoendog, statsmodels now knows which variable to predict, aligning with how linear regression is supposed to work. - Type Standardization: Converting everything to
float64ensures thatnp.isfinite(used to check for invalid values in the dataset) can process the matrix without throwing a type error.
After making these changes, you should see the OLS summary print out correctly, and you can proceed with your backward elimination process.
内容的提问来源于stack exchange,提问作者Anjali

