You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Andrew Ng决策边界非训练集属性的含义及线性回归相关疑问

Let me break this down clearly for you— I remember wrestling with this exact confusion when I first worked through Andrew Ng's course!

What "Decision boundaries aren't a property of the training set" really means

First, let's define what counts as a property of the training set: these are inherent, fixed characteristics of the raw data you're given. Examples include:

  • The specific feature values of each sample (e.g., "this patient’s tumor is 3cm wide")
  • The labels tied to those samples (e.g., "this tumor is malignant")
  • Statistical traits of the raw data (e.g., the average tumor size across all training samples)

A decision boundary, by contrast, is a product of the model you train using the training set—it doesn’t exist in the raw data itself. Here’s a more concrete way to see it:

  • Your training set is just a collection of points plotted on a graph. There’s no pre-drawn line or curve separating classes in the raw data.
  • When you run a classification algorithm (like logistic regression), it learns parameters that define the decision boundary. If you re-run the algorithm with a subset of the training data, or switch to a different model (like SVM), you’ll get a different boundary—even though the original training set’s core properties (the points themselves) haven’t changed.
  • The decision boundary is how your model interprets the data to make predictions, not a feature of the data itself.
Does this apply to linear regression's fitted lines/curves?

Absolutely—Andrew Ng’s point extends directly here. Let’s apply the same logic:

  • The training set for linear regression is a collection of (input, output) pairs (e.g., "1000 sq ft house sold for $200k"). These pairs are the inherent properties of the data.
  • The fitted line (or curve, for polynomial regression) is what the algorithm computes to minimize error between its predictions and training labels. It’s a model’s approximation of the relationship in the data, not something present in the raw training set.
  • If you remove a few outliers from the training set, or use a different regression technique (like LASSO instead of ordinary least squares), the fitted line will shift—yet the original training data’s core properties (the original (x,y) pairs) remain unchanged.
A simple analogy to drive this home

Think of your training set as a pile of baking ingredients (flour, eggs, sugar). The decision boundary/fitted line is the cake you bake using those ingredients. The ingredients have fixed properties (flour is powdery, eggs are liquid), but the cake’s shape, flavor, and texture aren’t properties of the ingredients—they’re the result of how you combine and process them. Different bakers (models/algorithms) could make very different cakes from the exact same ingredients.

内容的提问来源于stack exchange,提问作者cosmicShade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:59:05