关于Catboost二分类CTR任务类别变量编码默认设置的技术问询
Hey there! I've dug into CatBoost's default behavior for categorical encoding in binary CTR tasks since you couldn't track down these details in the official docs. Here's a clear breakdown:
CatBoost Default Categorical Encoding Configs for Binary CTR Tasks
Default Encoding Method
- For binary CTR tasks, CatBoost defaults to its proprietary CatBoost Encoding—a specialized variant of target encoding (mean target encoding) paired with ordered boosting. Unlike simple count or bucket encoding, this method leverages target variable statistics while using ordered data splits and grouping to mitigate overfitting, tailored specifically for gradient boosting workflows.
Default Prior Value
- The default prior value (global mean) is automatically calculated as the overall target mean of the training dataset. For CTR tasks, this translates to the click-through rate of the training set:
total_clicks / total_samples. You don’t need to set this manually—it’s derived directly from your training data.
Encoding Logic for Unlabeled Test Sets
- When handling test sets with no target values, CatBoost uses precomputed encoding mappings from the training phase:
- During training, each categorical level gets an encoded value that balances its local target statistics with the global prior.
- For inference, these pre-trained mappings are directly applied to test set categorical features. No target-related calculations are done on the test data, eliminating the risk of data leakage.
内容的提问来源于stack exchange,提问作者genericname
相关产品推荐
相关产品推荐

