关于Keras中ImageDataGenerator类及相关代码的两个技术问题
Hey there! Let's dive into your questions about this Keras image generator code from Deep Learning with Python:
from keras.preprocessing.image import ImageDataGenerator train_datagen = ImageDataGenerator(rescale=1./255) test_datagen = ImageDataGenerator(rescale=1./255) train_generator = train_datagen.flow_from_directory( train_dir, target_size=(150, 150), batch_size=20, class_mode='binary') validation_generator = test_datagen.flow_from_directory( validation_dir, target_size=(150, 150), batch_size=20, class_mode='binary')
问题1:为什么训练集和验证集都只用了rescale=1./255?
Great question! First off, image pixel values normally range from 0 to 255. Dividing by 255 scales all those values down to the 0-1 range, which is a standard preprocessing step for deep learning models. Here's why it matters:
- It keeps the model's weight updates stable—big pixel values can cause gradient explosions or vanishing gradients, which mess up training.
- It standardizes the input range across all images, so the model can focus on learning meaningful features instead of adjusting to varying value scales.
And crucially, we have to apply the exact same preprocessing to the validation set as the training set. If we didn't rescale the validation data, the input distribution would be totally different from what the model learned during training, making our validation metrics useless for judging how well the model will perform on real data.
问题2:class_mode='binary'到底是什么意思,什么时候用它?
class_mode='binary' tells the generator that we're working on a binary classification task—think things like "cat vs. dog" or "this image has a car vs. it doesn't". When you set this, the generator automatically converts your class labels into a binary 0/1 format. For this setup, you'll usually pair it with a sigmoid activation function in your model's final layer and use binary_crossentropy as the loss function.
If you're doing multi-class classification (like distinguishing between 5 different types of flowers), you'd switch this to class_mode='categorical' instead—this gives you one-hot encoded labels, which pairs with a softmax final layer and categorical_crossentropy loss. For regression tasks (predicting a continuous value from an image), you can set class_mode=None.
内容的提问来源于stack exchange,提问作者Hossein

