基于InceptionV3多标签分类预测值过小及阈值设置咨询
Hey there! Let's break down your two key issues with fine-tuning InceptionV3 for multi-class multi-label classification:
1. Why are my predictions all tiny values?
Your current setup has a few critical mismatches for multi-label tasks that are likely causing those near-zero predictions. Let's fix them:
a. Wrong class_mode in data generators
By default, flow_from_directory uses class_mode='categorical', which is for single-label tasks (each image belongs to exactly one class). For multi-label classification (images can have multiple classes), you need to set class_mode='binary' so the generator outputs binary labels (0/1) for each class—this matches your binary_crossentropy loss function.
Update your generator code:
train_generator = train_datagen.flow_from_directory( args.train_dir, target_size=(IM_WIDTH, IM_HEIGHT), batch_size=batch_size, class_mode='binary' # Add this line! ) validation_generator = test_datagen.flow_from_directory( args.val_dir, target_size=(IM_WIDTH, IM_HEIGHT), batch_size=batch_size, class_mode='binary' # Add this line! )
b. Dataset structure doesn't fit multi-label
flow_from_directory assumes each subfolder is a single class, and images in that folder only belong to that class. If you're doing multi-label (one image has multiple classes), this structure won't work. Instead, use flow_from_dataframe with a CSV file that maps each image to its multiple labels.
Example CSV format (train_labels.csv):
| image_path | cat | dog | bird | ... |
|---|---|---|---|---|
| train/img1.jpg | 1 | 0 | 1 | ... |
| train/img2.jpg | 0 | 1 | 0 | ... |
Then adjust your data loading:
import pandas as pd train_df = pd.read_csv('train_labels.csv') train_generator = train_datagen.flow_from_dataframe( dataframe=train_df, x_col='image_path', y_col=['cat', 'dog', 'bird'], # List all your class columns target_size=(IM_WIDTH, IM_HEIGHT), batch_size=batch_size, class_mode='raw' # For multi-label, use 'raw' mode )
c. Too few training epochs
NB_EPOCHS = 3 is way too low for transfer learning + fine-tuning. Models need time to adapt to your custom data. Try starting with 10-20 epochs for transfer learning, then 20-30 for fine-tuning, and adjust based on validation loss/accuracy.
d. Validation set checks
While the validation set isn't the direct cause of tiny predictions, double-check:
- It has the same number of classes as the training set
- Its images use the same preprocessing as training
- Labels are correctly annotated (no mismatches)
2. How to set a threshold for detecting classes?
In multi-label classification, each output value is a probability (0-1) representing how likely the class is present. Here's how to use thresholds effectively:
a. Basic thresholding (start with 0.5)
A default threshold of 0.5 works for many cases. Use it to filter which classes are considered "present":
import numpy as np # Assume you've preprocessed your image to match InceptionV3's requirements preprocessed_img = ... # Shape: (1, 299, 299, 3) predictions = model.predict(preprocessed_img)[0] # Get the prediction array for your image threshold = 0.5 # Get indices of classes where probability > threshold predicted_class_indices = np.where(predictions > threshold)[0] # Map indices to class names (use your generator's class_indices) class_names = list(train_generator.class_indices.keys()) detected_classes = [class_names[idx] for idx in predicted_class_indices] print("Detected classes in the image:", detected_classes)
b. Optimize the threshold for your task
0.5 isn't always optimal. Adjust based on what matters more:
- Reduce false positives (don't label a class that's not there): Raise the threshold (e.g., 0.6-0.7)
- Reduce false negatives (don't miss a class that is there): Lower the threshold (e.g., 0.3-0.4)
To find the best threshold, calculate the F1-score (balances precision and recall) on your validation set:
from sklearn.metrics import f1_score # Get validation set true labels and predictions val_true = validation_generator.labels # Or load from CSV if using flow_from_dataframe val_pred = model.predict(validation_generator) best_threshold = 0.5 best_f1 = 0 # Test thresholds from 0.1 to 0.9 in 0.1 steps for threshold in np.arange(0.1, 1.0, 0.1): # Convert probabilities to binary labels pred_labels = (val_pred > threshold).astype(int) # Calculate macro F1-score (treats all classes equally) current_f1 = f1_score(val_true, pred_labels, average='macro') if current_f1 > best_f1: best_f1 = current_f1 best_threshold = threshold print(f"Best threshold: {best_threshold} | Best F1-score: {best_f1}")
内容的提问来源于stack exchange,提问作者ou2105

