使用Pickle序列化H2O GBM模型遇__new__参数缺失错误求解决
First, let’s break down why you’re hitting that __new__() missing 1 required positional argument: 'keyvals' error: H2O model objects aren’t built to be serialized with pickle directly. They’re tightly tied to the H2O cluster’s internal state, and pickle only captures shallow references to cluster-side model data—not the actual weights, structure, or learned parameters. When you try to unpickle, the object can’t locate the underlying cluster resources it needs (the keyvals parameter refers to these cluster-side references).
Since you can’t use H2O’s official POJO/MOJO or default save methods, here are two workable alternatives:
Option 1: Serialize the Model’s JSON Representation
H2O models expose a JSON structure that contains all the model’s parameters, tree definitions, and learned weights. You can serialize this JSON instead of the model object itself, then reconstruct the model from the JSON when needed.
Code Example:
import h2o import pickle from h2o.estimators.gbm import H2OGradientBoostingEstimator # Initialize H2O and train model as before h2o.init() csv_url = "https://h2o-public-test-data.s3.amazonaws.com/smalldata/wisc/wisc-diag-breast-cancer-shuffled.csv" data = h2o.import_file(csv_url) y = 'diagnosis' x = data.columns[1:] # Simplified column selection train, test = data.split_frame(ratios=[0.75], seed=1) model = H2OGradientBoostingEstimator(distribution='bernoulli', ntrees=100, max_depth=4, learn_rate=0.1) model.train(x=x, y=y, training_frame=train, validation_frame=test) # Serialize the model's JSON instead of the model object model_json = model._model_json saved_model = pickle.dumps(model_json) # Later, to deserialize and reconstruct the model loaded_json = pickle.loads(saved_model) # Reinitialize H2O if your cluster was restarted h2o.init() # Create a new empty GBM estimator and load the JSON loaded_model = H2OGradientBoostingEstimator() loaded_model._model_json = loaded_json # Verify it works perf = loaded_model.model_performance(test) print(perf.auc())
Option 2: Save MOJO Bytes to Pickle (Indirect MOJO Use)
Even if you can’t use the standard MOJO deployment workflow, you can download the MOJO as a byte stream, serialize that with pickle, then load it back into H2O when needed. This works because the MOJO is a self-contained representation of the model.
Code Example:
import h2o import pickle from h2o.estimators.gbm import H2OGradientBoostingEstimator from h2o.utils.shared_utils import download_mojo # Train model as before h2o.init() csv_url = "https://h2o-public-test-data.s3.amazonaws.com/smalldata/wisc/wisc-diag-breast-cancer-shuffled.csv" data = h2o.import_file(csv_url) y = 'diagnosis' x = data.columns[1:] train, test = data.split_frame(ratios=[0.75], seed=1) model = H2OGradientBoostingEstimator(distribution='bernoulli', ntrees=100, max_depth=4, learn_rate=0.1) model.train(x=x, y=y, training_frame=train, validation_frame=test) # Download MOJO as bytes and serialize with pickle mojo_bytes = download_mojo(model, path=None, get_bytes=True) saved_model = pickle.dumps(mojo_bytes) # Later, deserialize and load the MOJO into H2O loaded_mojo_bytes = pickle.loads(saved_model) # Reinitialize H2O if needed h2o.init() # Load the MOJO into a new model object loaded_model = h2o.import_mojo_from_bytes(loaded_mojo_bytes) # Evaluate the model perf = loaded_model.model_performance(test) print(perf.auc())
Key Notes:
- Both methods avoid the cluster binding issue because they serialize actual model data, not just cluster references.
- Option 1 uses H2O’s internal JSON structure, which works best with the same H2O version (since JSON structure can change between releases).
- Option 2 uses the MOJO format, which is more portable across H2O versions and even supports offline prediction if needed.
内容的提问来源于stack exchange,提问作者Ankit Kothari

