使用sklearn中permutation_importance时出现样本数不匹配错误
解决IsolationForest局部特征重要性计算时的样本数不匹配问题
问题描述
编写Python代码为异常分数计算局部特征重要性时,触发以下ValueError:
ValueError: Found input variables with inconsistent numbers of samples: [1000, 1]
完整报错回溯:
Traceback (most recent call last): File "/tmp/sessions/827ab297d3bdaadf/main.py", line 44, in importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), y=np.zeros(df.shape[0]), scoring='neg_mean_squared_error') File "/usr/local/lib/python3.10/site-packages/sklearn/inspection/_permutation_importance.py", line 257, in permutation_importance baseline_score = _weights_scorer(scorer, estimator, X, y, sample_weight) File "/usr/local/lib/python3.10/site-packages/sklearn/inspection/_permutation_importance.py", line 19, in _weights_scorer return scorer(estimator, X, y) File "/usr/local/lib/python3.10/site-packages/sklearn/metrics/_scorer.py", line 219, in __call__ return self._score( File "/usr/local/lib/python3.10/site-packages/sklearn/metrics/_scorer.py", line 267, in _score return self._sign * self._score_func(y_true, y_pred, **self._kwargs) File "/usr/local/lib/python3.10/site-packages/sklearn/metrics/_regression.py", line 442, in mean_squared_error y_type, y_true, y_pred, multioutput = _check_reg_targets( File "/usr/local/lib/python3.10/site-packages/sklearn/metrics/_regression.py", line 100, in _check_reg_targets check_consistent_length(y_true, y_pred) File "/usr/local/lib/python3.10/site-packages/sklearn/utils/validation.py", line 387, in check_consistent_length raise ValueError( ValueError: Found input variables with inconsistent numbers of samples: [1000, 1]
相关代码:
import numpy as np import pandas as pd from sklearn.ensemble import IsolationForest from sklearn.inspection import permutation_importance # Set the random seed for reproducibility np.random.seed(42) # Generate the data num_sensors = 5 num_samples = 1000 data = np.random.randn(num_samples, num_sensors) for i in range(1, num_sensors): data[:, i] += data[:, i-1] # Add some anomalies to the data anomaly_indices = [100, 300, 500, 700] anomaly_magnitudes = [10, 5, 7, 12] for i, mag in zip(anomaly_indices, anomaly_magnitudes): data[i:i+5, 0] += mag # Convert the data to a pandas dataframe df = pd.DataFrame(data, columns=[f"sensor_{i}" for i in range(num_sensors)]) # Initialize the isolation forest model isolation_forest = IsolationForest(n_estimators=100, contamination='auto', random_state=42) # Fit the model to your data isolation_forest.fit(df) # Predict the anomalies anomalies = isolation_forest.predict(df) # Get the indices of anomalies anomaly_indices = np.where(anomalies == -1)[0] # Determine local feature importance for each anomaly for i in anomaly_indices: # Get the prediction score for the anomaly score = isolation_forest.score_samples(df.iloc[i].values.reshape(1, -1)) print(df.iloc[i].values.reshape(1, -1)) # Compute the permutation feature importance for the anomaly importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), y=np.zeros(df.shape[0]), scoring='neg_mean_squared_error') # Print the feature importance for the anomaly print(f"Anomaly detected in row {i}") print(f"Prediction score: {score}") print("Feature importance:") #for feature, importance_score in zip(df.columns, importance.importances_mean): #print(f"{feature}: {importance_score}") print("="*50)
报错原因分析
问题出在permutation_importance的参数传递上:
- 传入的
X是单个异常样本,形状为(1, 5)(仅1个样本) - 但
y参数传入了np.zeros(df.shape[0]),对应1000个样本的标签 - scikit-learn会强制校验
X和y的样本数必须一致,因此触发样本数不匹配的错误
此外,IsolationForest是无监督异常检测模型,本身不需要标签y,使用neg_mean_squared_error这类监督学习的评分函数也不符合无监督场景的逻辑。
修复方案
方案1:移除不必要的y参数,使用模型默认评分
直接去掉y参数,permutation_importance会自动使用IsolationForest的score_samples方法作为评分依据,这是无监督场景下的正确做法:
# 修改permutation_importance调用部分 importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), random_state=42)
方案2:自定义无监督评分函数(可选)
如果需要自定义评分逻辑,可以基于score_samples创建自定义scorer,比如:
from sklearn.metrics import make_scorer # 自定义评分函数:基于异常分数的负均值(适配permutation_importance的最大化逻辑) def anomaly_scorer(estimator, X): return -estimator.score_samples(X).mean() custom_scorer = make_scorer(anomaly_scorer) # 调用时使用自定义scorer,无需传入y importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), scoring=custom_scorer, random_state=42)
完整修复后的代码
import numpy as np import pandas as pd from sklearn.ensemble import IsolationForest from sklearn.inspection import permutation_importance # Set the random seed for reproducibility np.random.seed(42) # Generate the data num_sensors = 5 num_samples = 1000 data = np.random.randn(num_samples, num_sensors) for i in range(1, num_sensors): data[:, i] += data[:, i-1] # Add some anomalies to the data anomaly_indices = [100, 300, 500, 700] anomaly_magnitudes = [10, 5, 7, 12] for i, mag in zip(anomaly_indices, anomaly_magnitudes): data[i:i+5, 0] += mag # Convert the data to a pandas dataframe df = pd.DataFrame(data, columns=[f"sensor_{i}" for i in range(num_sensors)]) # Initialize the isolation forest model isolation_forest = IsolationForest(n_estimators=100, contamination='auto', random_state=42) # Fit the model to your data isolation_forest.fit(df) # Predict the anomalies anomalies = isolation_forest.predict(df) # Get the indices of anomalies anomaly_indices = np.where(anomalies == -1)[0] # Determine local feature importance for each anomaly for i in anomaly_indices: # Get the prediction score for the anomaly score = isolation_forest.score_samples(df.iloc[i].values.reshape(1, -1)) # Compute the permutation feature importance for the anomaly importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), random_state=42) # Print the feature importance for the anomaly print(f"Anomaly detected in row {i}") print(f"Prediction score: {score[0]:.4f}") print("Feature importance:") for feature, importance_score in zip(df.columns, importance.importances_mean): print(f"{feature}: {importance_score:.4f}") print("="*50)
内容的提问来源于stack exchange,提问作者pvalue0
相关产品推荐
相关产品推荐

