关于Andrew Ng课程中代价函数添加θ3、θ4正则项防过拟合的疑问
Great question—this gets to the core of how regularization works in machine learning, and why we don’t just tweak parameters by hand. Let’s break this down:
Machine learning is about learning parameters, not hardcoding them
The whole point of training a model is to let it find the optimal θ values from your data. If you manually force θ3 and θ4 to be small, you’re overriding the model’s ability to learn from the data. You might accidentally throw away useful signal from those features, or fail to balance their contribution properly relative to other features.Regularization adds a guided constraint, not a hard rule
The terms1000*(θ3)^2 + 1000*(θ4)^2act as a "penalty" in the cost function. When the model tries to make θ3 or θ4 very large (which would let those features dominate the prediction and cause overfitting), this penalty makes the total cost much higher. To minimize the overall cost, the algorithm automatically learns to keep θ3 and θ4 small—but not necessarily zero. It finds a balance between using the information in those features and avoiding overreliance on them.It integrates seamlessly with optimization algorithms
When you use gradient descent (or other optimization methods) to minimize the cost function, the regularization terms naturally modify the parameter update steps. For example, the update rule for θ3 would become:θ3 = θ3 - α * [ (1/m)∑(hθ(x(i)) - y(i))x3(i) + 2000*θ3 ]This extra
2000*θ3term pushes θ3 toward zero with each iteration, automatically. You don’t have to manually adjust anything—this is baked into the learning process.Manual adjustment is inflexible
If you manually set θ3 and θ4 to small values, you’re making a one-size-fits-all decision that might not adapt to different datasets or problem contexts. Regularization lets the model adapt: if the data actually needs θ3 to be a bit larger to make good predictions, the penalty will still keep it from getting too big, but won’t force it to be smaller than necessary.
内容的提问来源于stack exchange,提问作者Rahul

