预测商品销量是Regression还是Classification问题?机器学习新手求教
Understanding Why Sales Volume Prediction is a Regression Problem
Hey there, let's break this down clearly so you can see why Andrew Ng is framing this as a regression problem instead of classification.
First, let's recap the core difference between the two:
- Classification is about predicting categorical labels (e.g., "sold" vs "not sold", "cat" vs "dog") where outputs are distinct, non-quantifiable groups.
- Regression is about predicting quantifiable numerical values where the output has meaningful order or magnitude, and differences between values matter.
Your Misconception vs the Actual Problem
You framed the problem as "will the item be sold?" (a binary 0/1 classification), but the original question is predicting the exact number of units sold over 3 months—not just whether any units are sold. That's a key distinction.
Why "Discrete but Large-Range" Counts as Regression
You're right that sales volume is technically a discrete value (you can't sell half a unit), but here's why it's treated as regression:
- The values have quantifiable meaning: Selling 500 units vs 501 units is a meaningful difference (1 more unit sold), whereas classification labels like "category 1" vs "category 2" have no inherent numerical relationship.
- When the range of discrete values is large (e.g., 0 to 10,000 units), treating each possible number as a separate category for classification becomes impractical and misses the point. We care about how close our prediction is to the true sales number (e.g., being off by 10 units is better than being off by 1000), not whether we guessed the exact integer correctly. Regression models optimize for this kind of numerical error (like mean squared error), which aligns perfectly with sales prediction goals.
- In machine learning terminology, we often refer to these large-range discrete numerical targets as "continuous-like" because their predictive behavior aligns with regression tasks, not classification.
A Quick Example to Clarify
- If the question was: "Will this item sell more than 1000 units in 3 months?" That's a classification problem (yes/no).
- But the actual question is: "How many units will sell in 3 months?" That's regression, because we're predicting a numerical quantity, not a category.
Hopefully this clears up the confusion!
内容的提问来源于stack exchange,提问作者Muthuraja Palaniappan
相关产品推荐
相关产品推荐

