You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sklearn中permutation_importance时出现样本数不匹配错误

解决IsolationForest局部特征重要性计算时的样本数不匹配问题

问题描述

编写Python代码为异常分数计算局部特征重要性时,触发以下ValueError:

ValueError: Found input variables with inconsistent numbers of samples: [1000, 1]

完整报错回溯:

Traceback (most recent call last):
  File "/tmp/sessions/827ab297d3bdaadf/main.py", line 44, in 
    importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), y=np.zeros(df.shape[0]), scoring='neg_mean_squared_error')
  File "/usr/local/lib/python3.10/site-packages/sklearn/inspection/_permutation_importance.py", line 257, in permutation_importance
    baseline_score = _weights_scorer(scorer, estimator, X, y, sample_weight)
  File "/usr/local/lib/python3.10/site-packages/sklearn/inspection/_permutation_importance.py", line 19, in _weights_scorer
    return scorer(estimator, X, y)
  File "/usr/local/lib/python3.10/site-packages/sklearn/metrics/_scorer.py", line 219, in __call__
    return self._score(
  File "/usr/local/lib/python3.10/site-packages/sklearn/metrics/_scorer.py", line 267, in _score
    return self._sign * self._score_func(y_true, y_pred, **self._kwargs)
  File "/usr/local/lib/python3.10/site-packages/sklearn/metrics/_regression.py", line 442, in mean_squared_error
    y_type, y_true, y_pred, multioutput = _check_reg_targets(
  File "/usr/local/lib/python3.10/site-packages/sklearn/metrics/_regression.py", line 100, in _check_reg_targets
    check_consistent_length(y_true, y_pred)
  File "/usr/local/lib/python3.10/site-packages/sklearn/utils/validation.py", line 387, in check_consistent_length
    raise ValueError(
ValueError: Found input variables with inconsistent numbers of samples: [1000, 1]

相关代码:

import numpy as np
import pandas as pd
from sklearn.ensemble import IsolationForest
from sklearn.inspection import permutation_importance


# Set the random seed for reproducibility
np.random.seed(42)

# Generate the data
num_sensors = 5
num_samples = 1000
data = np.random.randn(num_samples, num_sensors)
for i in range(1, num_sensors):
    data[:, i] += data[:, i-1]

# Add some anomalies to the data
anomaly_indices = [100, 300, 500, 700]
anomaly_magnitudes = [10, 5, 7, 12]
for i, mag in zip(anomaly_indices, anomaly_magnitudes):
    data[i:i+5, 0] += mag

# Convert the data to a pandas dataframe
df = pd.DataFrame(data, columns=[f"sensor_{i}" for i in range(num_sensors)])

# Initialize the isolation forest model
isolation_forest = IsolationForest(n_estimators=100, contamination='auto', random_state=42)

# Fit the model to your data
isolation_forest.fit(df)

# Predict the anomalies
anomalies = isolation_forest.predict(df)

# Get the indices of anomalies
anomaly_indices = np.where(anomalies == -1)[0]

# Determine local feature importance for each anomaly
for i in anomaly_indices:
    # Get the prediction score for the anomaly
    score = isolation_forest.score_samples(df.iloc[i].values.reshape(1, -1))
    print(df.iloc[i].values.reshape(1, -1))
    # Compute the permutation feature importance for the anomaly
    importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), y=np.zeros(df.shape[0]), scoring='neg_mean_squared_error')

    # Print the feature importance for the anomaly
    print(f"Anomaly detected in row {i}")
    print(f"Prediction score: {score}")
    print("Feature importance:")
    #for feature, importance_score in zip(df.columns, importance.importances_mean):
        #print(f"{feature}: {importance_score}")
    print("="*50)

报错原因分析

问题出在permutation_importance的参数传递上:

  • 传入的X是单个异常样本,形状为(1, 5)(仅1个样本)
  • 但y参数传入了np.zeros(df.shape[0]),对应1000个样本的标签
  • scikit-learn会强制校验X和y的样本数必须一致,因此触发样本数不匹配的错误

此外,IsolationForest是无监督异常检测模型,本身不需要标签y,使用neg_mean_squared_error这类监督学习的评分函数也不符合无监督场景的逻辑。

修复方案

方案1:移除不必要的y参数,使用模型默认评分

直接去掉y参数,permutation_importance会自动使用IsolationForest的score_samples方法作为评分依据,这是无监督场景下的正确做法:

# 修改permutation_importance调用部分
importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), random_state=42)

方案2:自定义无监督评分函数(可选)

如果需要自定义评分逻辑,可以基于score_samples创建自定义scorer,比如:

from sklearn.metrics import make_scorer

# 自定义评分函数:基于异常分数的负均值(适配permutation_importance的最大化逻辑)
def anomaly_scorer(estimator, X):
    return -estimator.score_samples(X).mean()

custom_scorer = make_scorer(anomaly_scorer)

# 调用时使用自定义scorer,无需传入y
importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), scoring=custom_scorer, random_state=42)

完整修复后的代码

import numpy as np
import pandas as pd
from sklearn.ensemble import IsolationForest
from sklearn.inspection import permutation_importance


# Set the random seed for reproducibility
np.random.seed(42)

# Generate the data
num_sensors = 5
num_samples = 1000
data = np.random.randn(num_samples, num_sensors)
for i in range(1, num_sensors):
    data[:, i] += data[:, i-1]

# Add some anomalies to the data
anomaly_indices = [100, 300, 500, 700]
anomaly_magnitudes = [10, 5, 7, 12]
for i, mag in zip(anomaly_indices, anomaly_magnitudes):
    data[i:i+5, 0] += mag

# Convert the data to a pandas dataframe
df = pd.DataFrame(data, columns=[f"sensor_{i}" for i in range(num_sensors)])

# Initialize the isolation forest model
isolation_forest = IsolationForest(n_estimators=100, contamination='auto', random_state=42)

# Fit the model to your data
isolation_forest.fit(df)

# Predict the anomalies
anomalies = isolation_forest.predict(df)

# Get the indices of anomalies
anomaly_indices = np.where(anomalies == -1)[0]

# Determine local feature importance for each anomaly
for i in anomaly_indices:
    # Get the prediction score for the anomaly
    score = isolation_forest.score_samples(df.iloc[i].values.reshape(1, -1))
    # Compute the permutation feature importance for the anomaly
    importance = permutation_importance(isolation_forest, df.iloc[i].values.reshape(1, -1), random_state=42)

    # Print the feature importance for the anomaly
    print(f"Anomaly detected in row {i}")
    print(f"Prediction score: {score[0]:.4f}")
    print("Feature importance:")
    for feature, importance_score in zip(df.columns, importance.importances_mean):
        print(f"{feature}: {importance_score:.4f}")
    print("="*50)

内容的提问来源于stack exchange,提问作者pvalue0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 14:18:18