使用Sklearn与Keras Wrappers时pipeline.fit()运行失败的问题
Alright, let's break down these two errors you're facing when combining Scikit-learn Pipeline with Keras wrappers, and get you the trainable history object you're expecting.
Error 1: ValueError: not enough values to unpack (expected 2, got 1)
This boils down to a mismatch between Keras' wrapper API and Scikit-learn's expectations. The native KerasClassifier returns a History object when you call fit()—but Scikit-learn's Pipeline requires every step's fit() method to return the estimator itself (i.e., self). When Pipeline hits the Keras step, it expects to get back the estimator instance to proceed, but instead gets the History object, which causes that unpacking error.
Also, a quick side note: if you pass verbose=1 directly to pipeline.fit(), that parameter gets passed to all steps in the pipeline. The StandardScaler doesn't accept a verbose argument, but that usually throws a "unexpected keyword argument" error—so the main culprit here is definitely the non-standard return value from KerasClassifier.fit().
Error 2: AttributeError: 'numpy.ndarray' object has no attribute 'fit'
This means one of the steps in your Pipeline is a numpy array instead of a valid Scikit-learn-compatible estimator. The most common cause is:
- You're adding a raw Keras
Sequentialmodel directly to the Pipeline, instead of wrapping it withKerasClassifier. For example:# Wrong: Directly using Sequential without wrapper Pipeline([('scaler', StandardScaler()), ('model', Sequential())]) - Or your
build_fn(the function that creates your Keras model) is accidentally returning a numpy array instead of a compiled Keras model. Double-check that function to make sure it ends withreturn model.
Fixes & Workarounds
Let's get this working properly, including saving/loading the training history.
1. Make KerasClassifier Scikit-learn Compliant
We'll create a custom subclass of KerasClassifier that returns self from fit() (as Scikit-learn expects) while still storing the History object as an attribute:
from tensorflow.keras.wrappers.scikit_learn import KerasClassifier class SklearnFriendlyKerasClassifier(KerasClassifier): def fit(self, x, y, **kwargs): # Run the original fit and save the history self.history = super().fit(x, y, **kwargs) # Return self to match Scikit-learn's API return self
Use this subclass instead of the native KerasClassifier in your Pipeline.
2. Build Your Pipeline Correctly
Ensure your model step uses the wrapper, and your build_fn properly returns a compiled Keras model:
from sklearn.pipeline import Pipeline from sklearn.preprocessing import StandardScaler from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense def build_keras_model(): model = Sequential([ Dense(32, activation='relu', input_shape=(your_input_dim,)), Dense(1, activation='sigmoid') # Adjust based on your task ]) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) return model # Correct Pipeline setup self.pipeline = Pipeline([ ('scaler', StandardScaler()), ('model', SklearnFriendlyKerasClassifier(build_fn=build_keras_model, epochs=10, batch_size=32)) ])
3. Get & Save the History Object
Once you run fit(), you can access the History object through the Pipeline's named steps. To save it for later, serialize the history dictionary inside the History object (the History object itself can be tricky to pickle directly):
import pickle # Fit the pipeline self.pipeline.fit(self.X, self.Y) # Grab the training history training_history = self.pipeline.named_steps['model'].history # Save the history dictionary with open('model_history.pkl', 'wb') as f: pickle.dump(training_history.history, f) # Load it later with open('model_history.pkl', 'rb') as f: loaded_history = pickle.load(f) # Use loaded_history to plot loss/accuracy, etc.
4. Pass Fit Parameters Correctly
If you need to pass parameters like verbose or validation_data to Keras' fit(), use Scikit-learn's parameter naming convention: prefix the parameter with your model step name plus two underscores (model__). This ensures the parameter only goes to the Keras step, not the scaler:
# Correct way to pass verbose and validation data self.pipeline.fit(self.X, self.Y, model__verbose=1, model__validation_data=(X_val, Y_val))
Wrap-Up
- Fix Error 1 by using the custom
SklearnFriendlyKerasClassifierthat adheres to Scikit-learn's API. - Fix Error 2 by ensuring your Pipeline uses the wrapped Keras model and your
build_fnreturns a valid model. - Access and save the history via
pipeline.named_steps['model'].historyfor later use.
内容的提问来源于stack exchange,提问作者codebrotherone

