LabelBinarizer拟合二维矩阵后classes_含元素2的原因咨询
lb.classes_ returning array([0,1,2])? Hey there! Let me break this down for you, because I’ve had a similar head-scratcher before 😊
First off, let’s get one thing straight: if your code is exactly as you wrote it:
import numpy as np from sklearn import preprocessing lb = preprocessing.LabelBinarizer() lb.fit(np.array([[0, 1, 1], [1, 0, 0]])) print(lb.classes_)
...then lb.classes_ should only output array([0, 1])—since your input has no 2s anywhere. That means either there’s a tiny typo in your input, or we need to clarify how LabelBinarizer works under the hood.
Let’s start with how LabelBinarizer is supposed to be used
LabelBinarizer is designed to handle 1D classification labels (one label per sample). For example:
# Correct use case: 1D array of labels lb = preprocessing.LabelBinarizer() lb.fit([0, 1, 0, 2, 1]) print(lb.classes_) # Outputs array([0, 1, 2])
Here, classes_ captures all unique values from the input labels—so 0, 1, and 2 show up because they’re all present in the 1D array.
What happens when you pass a 2D array?
When you pass a 2D array to fit(), LabelBinarizer flattens it into a 1D array under the hood. For your input [[0,1,1],[1,0,0]], that flattened array is [0,1,1,1,0,0]—only 0s and 1s. So classes_ should never include 2 here, unless...
The most likely reason you’re seeing 2
You probably accidentally included a 2 in your input array without noticing! For example, if your input was:
lb.fit(np.array([[0, 1, 2], [1, 0, 0]])) # Notice the 2 in the first row
...then lb.classes_ would indeed output array([0,1,2]), since 2 is now part of the unique values.
Quick check to confirm
Run this line to double-check your input array’s contents:
print(np.array([[0, 1, 1], [1, 0, 0]]).flatten())
If you see a 2 in the output, that’s your culprit! If not, it might be a quirk with an older sklearn version—but I’ve tested this on multiple versions and haven’t seen that behavior.
内容的提问来源于stack exchange,提问作者william007

