如何获取启用早停的H2O GBM模型实际使用的树数量
Hey Dan, I totally get your frustration here—when you train an H2O GBM with ntrees=10000 and early stopping enabled, checking model.params['ntrees'] only gives you the upper limit you set, not the actual number of trees the model stopped at once early stopping kicked in. Good news, there are far more convenient ways than sifting through the score history!
Quickest Methods to Get the Actual Tree Count
Direct JSON Access: The fastest way is to pull the value straight from the model's internal JSON output:
model._model_json['output']['number_of_trees']This will return the exact number of trees the model ended up using (in your case, 280) without any extra steps.
Model Summary: If you prefer a more human-readable output, run
model.summary()—look for the "Number of Trees" line in the resulting model overview, it’ll show the actual count clearly.Pandas Shortcut (if you have pandas installed): If you’re already working with pandas, you can grab the last entry from the score history in one line:
model.score_history().iloc[-1]['number_of_trees']This is cleaner than scrolling through the full score history, though not as fast as the JSON method.
Why model.params['ntrees'] Isn’t the Right Value
As you noticed, model.params['ntrees'] returns {'default': 50, 'actual': 10000}—that actual value is just the maximum number of trees you allowed the model to train, not how many it actually used before early stopping triggered. The model stops once it hits the early stopping criteria, so this parameter only reflects your upper limit, not the final model size.
内容的提问来源于stack exchange,提问作者Dan

