子类化sklearn LinearSVC适配GridSearchCV时遇数值错误排查
Let's cut to the chase: your ValueError about NaN/inf values isn't actually a problem with your data—it's a missing return statement in the methods you overrode in your LinearSVCSub class.
Here's the core issue: sklearn's estimator API requires methods like fit, predict, score, and decision_function to return specific values (e.g., fit returns the estimator instance itself, predict returns the predicted labels). In your code, you call the parent class's method but don't pass its result back to the caller. For example:
def predict(self, X): X = self.transform_this(X) super(LinearSVCSub, self).predict(X) # No return here!
This means the method returns None by default. When GridSearchCV tries to use these methods during cross-validation, it gets None instead of valid predictions/scores, which causes downstream errors that manifest as the NaN/inf message you're seeing.
Corrected Code
Here's the fixed version of your subclass, with the missing return statements added:
from sklearn.datasets import load_breast_cancer from sklearn.svm import LinearSVC from sklearn.model_selection import GridSearchCV RANDOM_STATE = 123 class LinearSVCSub(LinearSVC): def __init__(self, penalty='l2', loss='squared_hinge', additional_parameter1=1, additional_parameter2=100, dual=True, tol=0.0001, C=1.0, multi_class='ovr', fit_intercept=True, intercept_scaling=1, class_weight=None, verbose=0, random_state=None, max_iter=1000): super(LinearSVCSub, self).__init__(penalty=penalty, loss=loss, dual=dual, tol=tol, C=C, multi_class=multi_class, fit_intercept=fit_intercept, intercept_scaling=intercept_scaling, class_weight=class_weight, verbose=verbose, random_state=random_state, max_iter=max_iter) self.additional_parameter1 = additional_parameter1 self.additional_parameter2 = additional_parameter2 def fit(self, X, y, sample_weight=None): X = self.transform_this(X) return super(LinearSVCSub, self).fit(X, y, sample_weight) # Added return def predict(self, X): X = self.transform_this(X) return super(LinearSVCSub, self).predict(X) # Added return def score(self, X, y, sample_weight=None): X = self.transform_this(X) return super(LinearSVCSub, self).score(X, y, sample_weight) # Added return def decision_function(self, X): X = self.transform_this(X) return super(LinearSVCSub, self).decision_function(X) # Added return def transform_this(self, X): return X if __name__ == '__main__': data = load_breast_cancer() X, y = data.data, data.target # Parameter tuning with custom LinearSVC param_grid = {'C': [0.00001, 0.0001, 0.0005], 'dual': (True, False), 'random_state': [RANDOM_STATE], 'additional_parameter1': [0.90, 0.80, 0.60, 0.30], 'additional_parameter2': [20, 30]} gs_model = GridSearchCV(estimator=LinearSVCSub(), verbose=1, param_grid=param_grid, scoring='roc_auc', n_jobs=-1) gs_model.fit(X, y)
Key Notes
- Always stick to sklearn's estimator API when subclassing: every method that returns something in the parent class must return the same type of value in your subclass. This ensures compatibility with tools like GridSearchCV and Pipeline.
- Your
transform_thismethod is currently a no-op, but when you expand its functionality later, just make sure it returns a valid feature matrix (no NaN, infinity, or values outside thefloat64range) and the rest of the pipeline will work as expected.
内容的提问来源于stack exchange,提问作者Hawklaz

