基于多层感知器(MLP)的图像分类:新增相似图像为何提升性能?
Great question—this cuts right to a foundational concept in machine learning: the gap between fitting a single example and generalizing to new, unseen data. Let’s break this down step by step:
Single-image training leads to overfitting, not learning patterns
When you train an MLP on just one image, backpropagation does work—it adjusts weights and biases to drive the prediction error for that specific image to nearly zero. But what the model learns isn’t the underlying concept (like "this is a cat" or "this is a handwritten 7"). Instead, it’s memorizing the exact pixel values and noise of that one image. Think of it as cramming for a test by memorizing one question and answer—you’ll nail that question, but fail every other similar one.More similar images reveal shared, meaningful features
Even "similar" images have subtle variations: different angles, lighting, backgrounds, or minor details (like a cat’s ear tilted slightly, or a digit written with a thicker stroke). These variations force the model to look past individual pixel quirks and focus on the common traits that define the category. For example, training on 100 cat images will push the model to learn that "pointed ears, whiskers, and a furry texture" are consistent markers of a cat, not just the specific white patch on one cat’s chest.Multiple samples stabilize weight updates
The gradient calculated from a single image is noisy—random pixel artifacts or outliers can skew the direction of weight updates. When you use a batch of images, you average the gradients across all samples. This gives you a more reliable, statistically sound direction for adjusting weights, moving the model closer to a set of parameters that work well for the entire data distribution, not just one outlier.The goal is generalization, not perfect training accuracy
Machine learning isn’t about getting 100% accuracy on the training data—it’s about performing well on data the model has never seen before. A model trained on one image will have near-perfect training accuracy but terrible generalization. Adding more similar images expands the model’s exposure to the range of inputs it might encounter in the real world, making its predictions more robust.
To put it simply: one image lets the model "memorize" a single case, while many similar images teach it to "understand" the rule.
内容的提问来源于stack exchange,提问作者Sung-IL

