如何用Scala从Sparkling Water导出H2O模型为MOJO并离线加载
Got it, let's tackle this since the official docs are a bit sparse on Scala examples for this exact workflow. Below is a complete, step-by-step breakdown for both exporting your Sparkling Water-trained model and loading it in a standalone Scala app without Spark or H2O clusters.
Step 1: Export the Model from Sparkling Water (Scala)
You're on the right track using ModelSerializationSupport—this is the correct utility to export a MOJO (Model Object, Optimized) which is lightweight and perfect for offline inference. Here's a full code snippet that includes model training and export:
import org.apache.spark.h2o._ import org.apache.spark.sql.SparkSession import hex.gbm.GBM import hex.gbm.GBMModel.GBMParameters import water.support.ModelSerializationSupport // Initialize Spark and Sparkling Water val spark = SparkSession.builder() .appName("SparklingWaterModelExport") .master("local[*]") // Adjust for your cluster setup .getOrCreate() val h2oContext = H2OContext.getOrCreate(spark) // Load sample data (replace with your dataset) val df = spark.read.option("header", "true").csv("path/to/your/data.csv") val h2oFrame = h2oContext.asH2OFrame(df) h2oFrame.replace(col = "label", h2oFrame.vec("label").toCategoricalVec()) // Ensure target is categorical if needed // Define GBM parameters (adjust based on your model) val gbmParams = new GBMParameters() gbmParams.train = h2oFrame gbmParams.response_column = "label" gbmParams.ntrees = 50 gbmParams.max_depth = 5 // Train the model val gbm = new GBM(gbmParams) val gbmModel = gbm.trainModel.get // Export the MOJO to a directory (this will create a zip file) val exportPath = "/path/to/export/model.mojo" ModelSerializationSupport.exportMOJO(gbmModel, exportPath, true) // Cleanup h2oContext.stop(true) spark.stop()
Key notes here:
- We export a MOJO instead of a POJO because MOJOs are self-contained, optimized, and work seamlessly with
hex-genmodelwithout needing Spark/H2O runtime. - The third parameter
trueinexportMOJOensures we zip the MOJO files into a single archive, which is easier to handle for deployment.
Step 2: Import the MOJO in a Spark-Free Scala App
Now, in your standalone Scala application (no Spark/H2O cluster required), you'll use the hex-genmodel library to load the MOJO and run predictions.
First, add the dependency
If using SBT, add this to your build.sbt:
libraryDependencies += "ai.h2o" % "hex-genmodel" % "3.46.0.3" // Match your Sparkling Water/H2O version
Then, load the model and run predictions
Here's the code to load the MOJO and make predictions on new data:
import hex.genmodel.easy.EasyPredictModelWrapper import hex.genmodel.easy.RowData import hex.genmodel.MojoModel // Load the MOJO zip file val mojoModel = MojoModel.load("/path/to/export/model.mojo.zip") // Wrap the model for easy prediction val easyModel = new EasyPredictModelWrapper(mojoModel) // Create a sample input row (match your model's feature schema) val row = new RowData() row.put("feature1", 1.2) row.put("feature2", "category_value") row.put("feature3", 45) // Run prediction val prediction = easyModel.predict(row) // Access prediction results (adjust based on your model type: classification/regression) println(s"Predicted class: ${prediction.classPrediction}") println(s"Class probabilities: ${prediction.classProbabilities.mkString(", ")}")
Important Notes
- Version Matching: Ensure the
hex-genmodelversion matches the H2O/Sparkling Water version used to train the model. Mismatched versions can cause compatibility issues. - Feature Schema: The input
RowDatamust exactly match the feature names and types used during model training. Any missing or mismatched features will throw errors. - No Runtime Dependencies: The
hex-genmodellibrary is a lightweight jar that doesn't require Spark or H2O cluster dependencies—perfect for embedding in microservices or standalone apps.
内容的提问来源于stack exchange,提问作者gerben

