CatBoost报错:无法计算需对数概率的绝对值相关指标
Let's break down what's causing this error and how to fix it quickly.
The Root Cause
Your error stems from passing use_best_model=True (via cat.get_params()) to the cv() function. This parameter is designed for single-model training—where you have a fixed train/validation split to pick the best iteration—but it doesn't make sense in cross-validation. CatBoost handles fold-specific validation internally, so keeping this parameter triggers a mismatch that breaks metric calculation when using eval_metric='Accuracy'.
Solution 1: Remove use_best_model from CV Parameters
Extract your model parameters, remove the incompatible use_best_model key, then pass the cleaned params to cv():
from catboost import Pool, CatBoostClassifier, cv from sklearn.model_selection import train_test_split from sklearn.metrics import accuracy_score import numpy as np # Split data train = data[:split] test = data[split:] # Get variables for a model x = train.drop(["Survived"], axis=1) y = train["Survived"] # Do train data splitting X_train, X_test, y_train, y_test = train_test_split(x,y, test_size=0.2, random_state=42) cat_features = np.where(x.dtypes != float)[0] # Initialize model with use_best_model (fine for fit(), not for cv()) cat = CatBoostClassifier(one_hot_max_size=7, iterations=21, random_seed=42, use_best_model=True, eval_metric='Accuracy') cat.fit(X_train, y_train, cat_features = cat_features, eval_set=(X_test, y_test)) pred = cat.predict(X_test) # Prepare pool and cleaned params for CV pool = Pool(X_train, y_train, cat_features=cat_features) cv_params = cat.get_params() cv_params.pop('use_best_model') # Remove the incompatible parameter cv_scores = cv(pool, cv_params, fold_count=10, plot=True) print('CV score: {:.5f}'.format(cv_scores['test-Accuracy-mean'].values[-1])) print('The test accuracy is :{:.6f}'.format(accuracy_score(y_test, cat.predict(X_test))))
Solution 2: Move use_best_model to the fit() Call
Alternatively, only specify use_best_model when calling fit() (since that's where it's actually needed), rather than including it in the model's initial parameters:
from catboost import Pool, CatBoostClassifier, cv from sklearn.model_selection import train_test_split from sklearn.metrics import accuracy_score import numpy as np # Split data train = data[:split] test = data[split:] # Get variables for a model x = train.drop(["Survived"], axis=1) y = train["Survived"] # Do train data splitting X_train, X_test, y_train, y_test = train_test_split(x,y, test_size=0.2, random_state=42) cat_features = np.where(x.dtypes != float)[0] # Initialize model WITHOUT use_best_model cat = CatBoostClassifier(one_hot_max_size=7, iterations=21, random_seed=42, eval_metric='Accuracy') # Pass use_best_model only to fit() cat.fit(X_train, y_train, cat_features = cat_features, eval_set=(X_test, y_test), use_best_model=True) pred = cat.predict(X_test) pool = Pool(X_train, y_train, cat_features=cat_features) # Now cat.get_params() won't include use_best_model, so cv() works cv_scores = cv(pool, cat.get_params(), fold_count=10, plot=True) print('CV score: {:.5f}'.format(cv_scores['test-Accuracy-mean'].values[-1])) print('The test accuracy is :{:.6f}'.format(accuracy_score(y_test, cat.predict(X_test))))
Both approaches resolve the conflict by ensuring use_best_model isn't passed to the cross-validation function, where it doesn't belong. The second option is cleaner since it keeps parameters scoped to where they're used.
内容的提问来源于stack exchange,提问作者Stanislav Jirák

