处理.jpg文件遇RuntimeError(异常20):InceptionV3重训练问题求助
Hey there, let's tackle your two main issues step by step, tailored to your setup (Ubuntu 16.04 LTS, GTX 1060, TensorFlow-GPU 1.6.0):
The core problem here is that the retrain.py you're using is built for the legacy TensorFlow InceptionV3 implementation, while your pre-trained model was trained using TensorFlow Slim's version of InceptionV3. These two versions have different variable naming conventions and slight structural differences, which causes the compatibility mismatch.
Here's how to fix it:
Swap in the Slim InceptionV3 definition in retrain.py:
- At the top of
retrain.py, add the import for Slim's InceptionV3:from tensorflow.contrib.slim.nets import inception import tensorflow.contrib.slim as slim - Replace the default model-building code with Slim's InceptionV3. Look for the section where the base model is defined (usually labeled something like "Create the model structure") and replace it with:
with slim.arg_scope(inception.inception_v3_arg_scope()): logits, end_points = inception.inception_v3( resized_input_tensor, num_classes=num_classes, is_training=is_training) - Adjust the checkpoint loading to match Slim's variable names. When loading your pre-trained MSCeleb-1M checkpoint, use
slim.assign_from_checkpoint_fninstead of the default checkpoint loader. Add this code before the training loop:checkpoint_path = "/path/to/your/mscelab-1m/inceptionv3/checkpoint" variables_to_restore = slim.get_variables_to_restore(exclude=['InceptionV3/Logits', 'InceptionV3/AuxLogits']) init_fn = slim.assign_from_checkpoint_fn(checkpoint_path, variables_to_restore) init_fn(sess)
This ensures you only restore the base model weights and leave the final classification layers to be retrained.
- At the top of
Alternative: Export your Slim model as a SavedModel:
If modifying the script feels too involved, you can export your pre-trained Slim InceptionV3 to a SavedModel format first, then updateretrain.pyto load this SavedModel instead of the legacy checkpoint. This avoids variable name mismatches entirely.
Error 20 in TensorFlow's JPG decoder usually points to one of two issues: corrupted JPG files or incompatible JPG formats (like CMYK instead of RGB), or a minor bug in TF 1.6.0's decoding logic.
Try these fixes in order:
- Audit your JPG files:
Some of your custom images might be corrupted or use non-standard formats. Run this quick script to identify and fix/remove bad files:from PIL import Image import os def validate_jpgs(image_dir): for root, _, files in os.walk(image_dir): for filename in files: if filename.lower().endswith('.jpg') or filename.lower().endswith('.jpeg'): file_path = os.path.join(root, filename) try: with Image.open(file_path) as img: img.verify() # Checks for corruption # Convert CMYK images to RGB if needed if img.mode != 'RGB': print(f"Converting CMYK to RGB: {file_path}") rgb_img = img.convert('RGB') rgb_img.save(file_path) except (IOError, SyntaxError) as e: print(f"Deleting bad file: {file_path}") os.remove(file_path) validate_jpgs("/path/to/your/custom/images") - Tweak the JPG decoding in retrain.py:
TF 1.6.0'stf.image.decode_jpeghas a known issue with certain JPGs. Modify the decoding line to use a more robust DCT method:
Find this code inretrain.py:
Replace it with:decoded_image = tf.image.decode_jpeg(image_data, channels=3)decoded_image = tf.image.decode_jpeg(image_data, channels=3, dct_method='INTEGER_ACCURATE') - Verify CUDA/cuDNN compatibility:
TensorFlow-GPU 1.6.0 requires CUDA 8.0 and cuDNN 6.0. Mismatched versions can cause random runtime errors, so double-check that your system has these exact versions installed.
Once you implement these changes, the retraining script should work smoothly with your custom dataset and the InceptionV3 model you trained on MSCeleb-1M.
内容的提问来源于stack exchange,提问作者Stephen Chai

