如何保存与加载sklearn中的FeatureUnion特征融合对象?
I’ve dealt with exactly this scenario before when working on RL tasks with sklearn-based feature transformers—saving and reusing FeatureUnion objects is straightforward once you know the right tools. Sklearn officially recommends using joblib for this purpose, as it’s optimized for objects with large numerical arrays (like the ones in your RBFSampler components) and outperforms the standard pickle module in both speed and storage efficiency.
Step 1: Install Joblib (if needed)
First, make sure you have joblib installed—it’s usually included with sklearn, but if not, run:
pip install joblib
Step 2: Save Your FeatureUnion and Scaler
Since your feature pipeline depends on both the fitted FeatureUnion and the fitted scaler, you need to save both to ensure consistent results later. You can save them as separate files or bundle them into a single dictionary for convenience:
Option A: Save as separate files
import joblib # Save the fitted featurizer joblib.dump(featurizer, 'rbf_featurizer.joblib') # Save the fitted scaler (critical—don't skip this!) joblib.dump(scaler, 'state_scaler.joblib')
Option B: Bundle into a single file
import joblib # Create a dictionary to hold both components preprocessing_pipeline = { 'featurizer': featurizer, 'scaler': scaler } # Save the bundle joblib.dump(preprocessing_pipeline, 'rbf_preprocessing_bundle.joblib')
Step 3: Load and Reuse the Saved Objects
When you need to reuse the pipeline later, load the saved files and reconstruct your feature function:
For separate files:
import joblib # Load the saved components loaded_featurizer = joblib.load('rbf_featurizer.joblib') loaded_scaler = joblib.load('state_scaler.joblib') # Recreate your feature function def featurise_state(state): scaled = loaded_scaler.transform([state]) featurised = loaded_featurizer.transform(scaled) return featurised[0]
For the bundled file:
import joblib # Load the bundle loaded_pipeline = joblib.load('rbf_preprocessing_bundle.joblib') # Extract components loaded_featurizer = loaded_pipeline['featurizer'] loaded_scaler = loaded_pipeline['scaler'] # Recreate the feature function (same as above) def featurise_state(state): scaled = loaded_scaler.transform([state]) featurised = loaded_featurizer.transform(scaled) return featurised[0]
Key Notes to Avoid Issues
- Keep sklearn versions consistent: Make sure the sklearn version you use to load the objects matches (or is compatible with) the version you used to save them. Breaking changes in serialization can cause errors if versions differ too much.
- Never refit the scaler: Always use the fitted scaler you saved—refitting on new data will change the mean/variance values, leading to mismatched features that break your RL model’s performance.
- Test the loaded pipeline: After loading, run a quick test with a sample state to verify that the output matches the original featurizer’s output. This ensures nothing went wrong during serialization.
内容的提问来源于stack exchange,提问作者EmptySet

