深度学习新手求教:CNN识别浦那贫民窟卫星图的图像尺寸问题
Hey Sunny, let's break down what might be going on with your slum detection model and walk through actionable fixes!
问题排查与解决方案
1. 先聊聊你怀疑的图像尺寸问题
Your hunch about image size is totally valid—satellite images rely heavily on fine-grained details, and squeezing them down to 128x128 might be gutting the features your model needs to distinguish slums:
- Critical feature loss: Slums are defined by dense, small structures, cluttered layouts, and unique roof patterns. At 128x128, these details get blurred to the point of invisibility, especially if you're using a naive scaling method like nearest-neighbor interpolation.
- Convolution kernel mismatch: A 3x3 kernel works great for most cases, but when your input is already tiny, repeated convolution + pooling layers will shrink feature maps to near-nothing. For example, 3 rounds of pooling would take 128→64→32→16—hardly enough space to capture meaningful patterns.
Quick fixes for size issues:
- Try scaling up your input to 256x256 or 512x512. Yes, it'll use more compute, but satellite image tasks almost always benefit from higher resolution. If hardware is tight, look into multi-scale training (feeding images at different sizes) or pyramid pooling layers to retain detail without full-size inputs.
- Swap to bilinear or bicubic interpolation when resizing images—these methods preserve edges better than nearest-neighbor scaling.
2. Training vs. input data distribution mismatch (the #1 culprit for "high train acc, bad predictions")
Even if size is fixed, a common trap with small datasets (100+100 is tiny for deep learning) is that your model might be memorizing training images instead of learning general slum patterns. Here's what to check:
- Are your input images from the same scale/region as training data? The Google Maps link you shared uses a specific zoom level (286m). If your test inputs are at a different zoom, the feature distribution will be totally off—slums look very different at 100m vs. 500m zoom.
- Is data augmentation missing? With only 200 total images, overfit is guaranteed unless you add variation. Try adding random rotations, flips, brightness/contrast adjustments, and random cropping to your training pipeline. This forces the model to learn universal slum traits, not just the exact pixels in your training set.
- Double-check label accuracy: Annotating satellite images is tricky—easy to mislabel edge areas or semi-slum regions. A few bad labels can throw off your model's learning.
3. Model & training pipeline checks
- Test for overfit: Compare training accuracy to a validation set (split your 200 images into 160 train /40 val). If train acc is near 100% but val acc is low, you need to add regularization: Dropout layers, L2 weight decay, or simplify your model architecture.
- Verify preprocessing consistency: Make sure your input images go through the exact same steps as training data—same normalization (mean/std scaling), same channel order (RGB vs BGR), same cropping/resizing logic. A tiny mismatch here can break predictions.
- Check output layer setup: For binary classification (slum vs non-slum), your output layer should use a
sigmoidactivation function withbinary cross-entropyloss. Mixing this up (e.g., using softmax for 2 classes) can lead to wonky predictions.
Quick validation steps to narrow down the issue
- Take a slum image from your training set, feed it directly into the model. If it classifies correctly, your model works—problem is with input data/preprocessing. If it fails, your training pipeline has issues (overfit, bad labels, etc.).
- Use tools like Grad-CAM to visualize what parts of the image your model is focusing on. If it's fixating on random noise instead of slum structures, you know your model isn't learning the right features.
内容的提问来源于stack exchange,提问作者SUNNY
相关产品推荐
相关产品推荐

