DiffPrivLib设ε=∞时决策树/随机森林与Sklearn结果不符求助
关于DiffPrivLib与Scikit-learn模型在epsilon=∞时结果不一致的问题
IBM专门设计了差分隐私库DiffPrivLib,使其使用方式与Scikit-learn完全一致以提升易用性。其官方教程指出,当设置epsilon=∞且使用相同random_state时,DiffPrivLib模型应与Scikit-learn的非私有模型完全一致。我已验证该特性在Gaussian Naive Bayes上有效,但使用几乎相同的代码框架运行Decision Tree Classifier时,得到的结果却截然不同;后续扩展到随机森林模型测试,也发现了同样的问题,求排查代码问题。
1. Scikit-learn非私有决策树代码及结果
# Import necessary packages import pandas as pd import numpy as np from sklearn.tree import DecisionTreeClassifier from sklearn.metrics import confusion_matrix, matthews_corrcoef from sklearn import datasets from sklearn.model_selection import train_test_split import diffprivlib as dp # Load breast cancer dataset into a dataframe dataset = datasets.load_breast_cancer() data = pd.DataFrame(data=dataset.data, columns=dataset.feature_names) data['target'] = dataset.target # Split into test and training sets X = data.drop(['target'], axis=1) y = data['target'] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2) # Build a Decision Tree Classifier dt = DecisionTreeClassifier(random_state=42) # Model training dt.fit(X_train, y_train) # Predict Output y_pred_dt = dt.predict(X_test) # Output metrics print(matthews_corrcoef(y_test, y_pred_dt)) print(confusion_matrix(y_test, y_pred_dt))
运行结果:
0.8170347321131809 [[40 4] [ 4 66]]
2. DiffPrivLib决策树(epsilon=∞)代码及结果
# Calculate the bounds to prevent a privacy warning with DP-RF min_values = data.min().tolist()[:-1] max_values = data.max().tolist()[:-1] bounds = (min_values, max_values) # Define the classes to prevent a privacy warning with DP-RF classes = np.array([0, 1]) # Build a Decision Tree Classifier DPdt = dp.models.DecisionTreeClassifier(random_state=42, epsilon=np.inf, bounds=bounds, classes=classes) # Model training DPdt.fit(X_train, y_train) # Predict Output y_pred_DPdt = DPdt.predict(X_test) # Output metrics print(matthews_corrcoef(y_test, y_pred_DPdt)) print(confusion_matrix(y_test, y_pred_DPdt))
运行结果:
0.2813017973520981 [[13 33] [ 5 63]]
3. 扩展测试:随机森林模型的问题
我用DiffPrivLib的Random Forest和sklearn乳腺癌数据集重新测试,发现存在同样问题:
# Load the required packages import pandas as pd import numpy as np import matplotlib.pyplot as plt from sklearn.metrics import confusion_matrix from sklearn.metrics import matthews_corrcoef from sklearn.ensemble import RandomForestClassifier from sklearn import datasets from sklearn.model_selection import train_test_split import diffprivlib as dp # Load breast cancer dataset into a dataframe dataset = datasets.load_breast_cancer() data = pd.DataFrame(data=dataset.data, columns=dataset.feature_names) data['target'] = dataset.target # Calculate the bounds to prevent a privacy warning with DP-RF min_values = data.min().tolist()[:-1] max_values = data.max().tolist()[:-1] bounds = (min_values, max_values) # Calculate the classes to prevent a privacy warning with DP-RF classes = np.array([0, 1]) # Split into test and training sets X = data.drop(['target'], axis=1) y = data['target'] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2) # Calculate the MCC for the non-DP version of RF rf = RandomForestClassifier(random_state=42) rf.fit(X_train, y_train) y_pred_rf = rf.predict(X_test) rf_MCC = matthews_corrcoef(y_test, y_pred_rf) print(confusion_matrix(y_test, y_pred_rf)) print("MCC: ", rf_MCC) # Build a Differentially Private Random Forest Classifier epsilon = np.inf # Example epsilon value for differential privacy DPrf = dp.models.RandomForestClassifier(epsilon=epsilon, random_state=42, bounds=bounds, classes=classes) DPrf.fit(X_train, y_train) y_pred_DPrf = DPrf.predict(X_test) y_pred_DPrf = y_pred_DPrf.astype(int) DPinf_MCC = matthews_corrcoef(y_test, y_pred_DPrf) # Output metrics print(confusion_matrix(y_test, y_pred_DPrf)) print("MCC: ", DPinf_MCC)
将测试在一系列epsilon值上重复30次并取平均值,得到如下结果:
可以看到DP版本的Random Forest收敛到的准确率远低于标准Scikit-learn版本。
请问我哪里出错了?
内容的提问来源于stack exchange,提问作者Alison Krauskopf
相关产品推荐
相关产品推荐

