Google Colab中sklearn.datasets.load_files无限运行问题求助
I’ve run into similar slowdowns with large datasets in Colab before, so here are targeted fixes to get your image loading back on track:
Fix path whitespace issues
The space inColab Notebooksmight be causing unexpected parsing behavior forload_files, even with quoted paths. Try either:- Renaming the directory to
ColabNotebooks(no spaces) and updating your path to"/content/cat_dog/ColabNotebooks/dataset/training_set" - Using a raw Python string to avoid escape character conflicts:
load_files(r"/content/cat_dog/Colab Notebooks/dataset/training_set")
- Renaming the directory to
Validate your dataset structure first
load_filesexpects the target directory to contain subfolders for each category (e.g.,training_set/cats/andtraining_set/dogs/). Confirm you’re pointing to the right place with these commands:# Check for category subfolders !ls "/content/cat_dog/Colab Notebooks/dataset/training_set" # Count total image files to match your 9000 count !find "/content/cat_dog/Colab Notebooks/dataset/training_set" -type f -name "*.jpg" | wc -lIf the count is off, you might be targeting the wrong directory, and
load_filescould be traversing unnecessary system files.Optimize load_files parameters
Default settings like shuffling and text decoding waste time on image data. Tweak these to speed things up:from sklearn.datasets import load_files dataset = load_files( "/content/cat_dog/Colab Notebooks/dataset/training_set", shuffle=False, # Disable shuffle (you can shuffle manually later) encoding=None, # Skip text encoding for image files decode_error="ignore" )Reset your Colab runtime
Sometimes Colab’s file system cache or runtime state gets corrupted, leading to slow IO. Go toRuntime > Restart runtimeand re-run your code from scratch.Switch to a GPU/TPU runtime
GPU/TPU runtimes in Colab have faster disk IO than CPU-only instances. Switch viaRuntime > Change runtime typeand select GPU as the hardware accelerator.Use an image-optimized loading method
load_filesisn’t built for large image datasets. For better performance, use TensorFlow/Keras’ImageDataGeneratorto load images in batches:from tensorflow.keras.preprocessing.image import ImageDataGenerator datagen = ImageDataGenerator(rescale=1./255) train_generator = datagen.flow_from_directory( "/content/cat_dog/Colab Notebooks/dataset/training_set", target_size=(150, 150), # Adjust to your image dimensions batch_size=32, class_mode='binary' # Use 'categorical' for more than 2 classes )
内容的提问来源于stack exchange,提问作者Jean Albert

