能否从H2O拟合的GLM模型中直接获取训练优化的惩罚似然?
Great question! I’ve run into this exact need before, and luckily H2O does expose the penalized likelihood value that’s optimized during GLM training—no manual recalculation required. Here’s how to get it:
Key Background
First, a quick clarification: h2o.logloss() and h2o.residual_deviance() return unpenalized metrics (they only capture the loglikelihood or deviance component). The penalized likelihood your model optimizes includes both the loglikelihood (or deviance for regression) plus the L1/L2 penalty terms you specified via lambda_ and alpha.
How to Extract the Penalized Likelihood
H2O stores this optimized value directly in the model’s training metrics. The exact method depends on whether you’re using Python or R:
In Python
After fitting your GLM model, access the training_metrics() object, then pull the penalized_loglikelihood field from its underlying JSON data:
# Assume your fitted model is named glm_model penalized_likelihood = glm_model.training_metrics()._metric_json['penalized_loglikelihood'] print(f"Final Penalized Log-Likelihood: {penalized_likelihood}")
In R
For R users, the value lives in the model’s training metrics slot:
# Assume your fitted model is named glm_model penalized_likelihood <- glm_model@model$training_metrics$penalized_loglikelihood cat("Final Penalized Log-Likelihood:", penalized_likelihood, "\n")
Example Workflow (Python)
Here’s a complete example to illustrate:
import h2o from h2o.estimators.glm import H2OGeneralizedLinearEstimator # Initialize H2O h2o.init() # Load sample classification data (built-in H2O dataset) data = h2o.load_dataset("prostate") data['CAPSULE'] = data['CAPSULE'].asfactor() # Fit a penalized GLM (binomial family with elastic net penalty) glm_model = H2OGeneralizedLinearEstimator( family="binomial", lambda_=0.1, # Regularization strength alpha=0.5 # Elastic net mix (0 = L2, 1 = L1) ) glm_model.train(x=['AGE', 'RACE', 'PSA', 'DCAPS'], y='CAPSULE', training_frame=data) # Extract and print the penalized likelihood penalized_ll = glm_model.training_metrics()._metric_json['penalized_loglikelihood'] print(f"Optimized Penalized Log-Likelihood: {penalized_ll:.4f}")
Verification Tip
If you want to confirm this value is correct, you can manually compute it as:
Penalized Likelihood = Loglikelihood + Penalty Term
You can pull the unpenalized loglikelihood from glm_model.training_metrics().loglikelihood() and calculate the penalty term using the model’s coefficients and regularization parameters. The result should match the penalized_loglikelihood value H2O provides.
内容的提问来源于stack exchange,提问作者James Hirschorn

