使用SelectKBest结合chi2做特征选择时报错ValueError: Input X must be non-negative求助
ValueError: Input X must be non-negative with SelectKBest & chi2 Hey there! I've run into this exact issue before, so let's break down what's happening and how to fix it.
Why the Error Happens
The chi-squared (chi2) feature selection method is built to work with non-negative feature data—it measures relationships between features and your target variable using frequency-based statistics, where negative values don't have meaningful statistical interpretation. That's why your code throws an error when x_train contains negative numbers.
Solutions to Try
Here are three practical fixes tailored to your code:
1. Scale Features to a Non-Negative Range
If you want to keep using chi2, the simplest fix is to scale your features to the [0, 1] interval with MinMaxScaler. This preserves relative relationships between data points while eliminating negatives. Update your code like this:
from sklearn.feature_selection import SelectKBest, chi2 from sklearn.model_selection import train_test_split from sklearn.preprocessing import MinMaxScaler # Original data preparation x = data.drop('TARGET_B', axis='columns') x = x.drop('TARGET_D', axis='columns') y = data.TARGET_B # Split train/test sets x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.20, random_state=None) # Scale features (fit only on training data to avoid data leakage!) scaler = MinMaxScaler() x_train_scaled = scaler.fit_transform(x_train) x_test_scaled = scaler.transform(x_test) # Now chi2 runs without errors select_feature = SelectKBest(chi2, k=15).fit(x_train_scaled, y_train)
2. Switch to a Method That Supports Negative Values
If scaling your data doesn't align with your use case (e.g., you need to preserve original feature distributions), use a feature selection method that tolerates negatives. For classification tasks like yours, great options are:
f_classif: Uses ANOVA F-values to measure feature importancemutual_info_classif: Measures mutual information between features and the target
Example with f_classif:
from sklearn.feature_selection import SelectKBest, f_classif # Keep your existing train/test split code x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.20, random_state=None) # Replace chi2 with f_classif select_feature = SelectKBest(f_classif, k=15).fit(x_train, y_train)
3. Clean Invalid Negative Values (If Applicable)
Sometimes negative values are anomalies (e.g., calculation errors leading to negative counts). If you're sure negatives don't belong in your dataset, clean them by replacing with a logical value like 0 or the feature's mean:
# Replace all negative values in x with 0 x = x.apply(lambda col: col.where(col >= 0, 0)) # Proceed with train/test split and feature selection as before
Note: Only do this if you've verified the negatives are invalid—don't arbitrarily modify data without understanding their origin!
内容的提问来源于stack exchange,提问作者Shah

