使用BayesSearchCV调参时遇unhashable type: 'dict'错误的解决方法
解决BayesSearchCV调优LogisticRegression时class_weight字典不可哈希的问题
问题原因
BayesSearchCV依赖skopt库处理参数空间,当你传入字典列表作为class_weight的候选值时,skopt会尝试将这些字典作为键构建映射表(用于类别型参数的编码),但字典是不可哈希的类型,无法作为字典的键,因此抛出TypeError: unhashable type: 'dict'。
解决方案
核心思路是:不直接传递字典作为参数候选,而是将字典中的权重值作为独立的搜索参数,在模型内部动态生成class_weight字典。以下是两种具体实现方式:
方法1:自定义包装器模型
创建一个继承自sklearn基础类的包装器,将权重值作为单独参数,内部自动构建class_weight字典:
import numpy as np from sklearn.base import BaseEstimator, ClassifierMixin from sklearn.linear_model import LogisticRegression from skopt import BayesSearchCV from skopt.space import Real # 自定义带权重参数的逻辑回归模型 class WeightedLogisticRegression(BaseEstimator, ClassifierMixin): def __init__(self, C=1.0, class_weight_x=1.0): self.C = C self.class_weight_x = class_weight_x # 动态构建class_weight字典 self.model = LogisticRegression(C=self.C, class_weight={0:1, 1:self.class_weight_x}) def fit(self, X, y): self.model.fit(X, y) return self def predict(self, X): return self.model.predict(X) def predict_proba(self, X): return self.model.predict_proba(X) def score(self, X, y, sample_weight=None): return self.model.score(X, y, sample_weight) # 定义搜索空间:直接搜索权重值而非字典 params = { 'C': Real(1.0, 5.0), 'class_weight_x': Real(1.0, 5.0) } # 初始化BayesSearchCV并训练 bs_cv = BayesSearchCV(WeightedLogisticRegression(), params, cv=50, n_iter=5, random_state=42, scoring='auprc', verbose=False) history = bs_cv.fit(X_train, y_train) best_bayes = bs_cv.best_estimator_ # 输出结果 print("Test set AUPRC: {}".format(bs_cv.score(X_test, y_test))) # 手动将参数转换为class_weight字典格式 print("Best parameters are: C={}, class_weight={{0:1, 1:{}}}".format(bs_cv.best_params_['C'], bs_cv.best_params_['class_weight_x']))
方法2:使用函数包装简化实现
如果不想写完整的包装器类,可以用函数动态生成模型,配合functools.partial让BayesSearchCV正确识别参数:
import numpy as np from sklearn.linear_model import LogisticRegression from skopt import BayesSearchCV from skopt.space import Real from functools import partial # 定义动态生成带权重逻辑回归模型的函数 def create_weighted_lr(C=1.0, class_weight_x=1.0): return LogisticRegression(C=C, class_weight={0:1, 1:class_weight_x}) # 用partial包装,确保参数能被搜索框架识别 weighted_lr = partial(create_weighted_lr) # 定义搜索空间 params = { 'C': Real(1.0, 5.0), 'class_weight_x': Real(1.0, 5.0) } # 初始化并训练 bs_cv = BayesSearchCV(weighted_lr, params, cv=50, n_iter=5, random_state=42, scoring='auprc', verbose=False) history = bs_cv.fit(X_train, y_train) best_bayes = bs_cv.best_estimator_ # 输出结果 print("Test set AUPRC: {}".format(bs_cv.score(X_test, y_test))) print("Best parameters are: C={}, class_weight={{0:1, 1:{}}}".format(bs_cv.best_params_['C'], bs_cv.best_params_['class_weight_x']))
注意事项
- 两种方法都将原本的字典参数拆分为可哈希的数值参数,彻底避免了skopt的哈希错误。
- 输出最优参数时,需要手动将
class_weight_x转换为字典形式,因为BayesSearchCV返回的是搜索的原始参数值,而非最终传入模型的class_weight字典。
内容的提问来源于stack exchange,提问作者guilerme_coppola84
相关产品推荐
相关产品推荐

