You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Andrew Ng课程中代价函数添加θ3、θ4正则项防过拟合的疑问

Why Add Regularization Terms Instead of Manually Reducing θ3 and θ4?

Great question—this gets to the core of how regularization works in machine learning, and why we don’t just tweak parameters by hand. Let’s break this down:

  • Machine learning is about learning parameters, not hardcoding them
    The whole point of training a model is to let it find the optimal θ values from your data. If you manually force θ3 and θ4 to be small, you’re overriding the model’s ability to learn from the data. You might accidentally throw away useful signal from those features, or fail to balance their contribution properly relative to other features.

  • Regularization adds a guided constraint, not a hard rule
    The terms 1000*(θ3)^2 + 1000*(θ4)^2 act as a "penalty" in the cost function. When the model tries to make θ3 or θ4 very large (which would let those features dominate the prediction and cause overfitting), this penalty makes the total cost much higher. To minimize the overall cost, the algorithm automatically learns to keep θ3 and θ4 small—but not necessarily zero. It finds a balance between using the information in those features and avoiding overreliance on them.

  • It integrates seamlessly with optimization algorithms
    When you use gradient descent (or other optimization methods) to minimize the cost function, the regularization terms naturally modify the parameter update steps. For example, the update rule for θ3 would become:

    θ3 = θ3 - α * [ (1/m)∑(hθ(x(i)) - y(i))x3(i) + 2000*θ3 ]
    

    This extra 2000*θ3 term pushes θ3 toward zero with each iteration, automatically. You don’t have to manually adjust anything—this is baked into the learning process.

  • Manual adjustment is inflexible
    If you manually set θ3 and θ4 to small values, you’re making a one-size-fits-all decision that might not adapt to different datasets or problem contexts. Regularization lets the model adapt: if the data actually needs θ3 to be a bit larger to make good predictions, the penalty will still keep it from getting too big, but won’t force it to be smaller than necessary.

内容的提问来源于stack exchange,提问作者Rahul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:13:32