You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

感知器训练规则中为何乘以x_i?对权重更新幅度的疑问

Understanding the Perceptron Update Rule's Dependence on Input Magnitude

Great question—this is a common point of confusion when first diving into perceptron training! Let’s break down the logic behind the rule from Tom Mitchell’s Machine Learning and address your intuition.

First, let’s restate the perceptron training rule clearly:
$$\Delta w_i = \eta (d_i - o_i) x_i$$

Where:

  • $\eta$: Training rate (controls the overall step size of all weight updates)
  • $d_i$: Expected output for the current training example
  • $o_i$: Actual output produced by the perceptron
  • $x_i$: The $i$-th input feature value

Why Larger $x_i$ Leads to Larger Weight Updates

Your observation about the weighted sum’s sensitivity to $w_i$ when $x_i$ is large is totally valid—but the update rule is designed to correct errors efficiently, not just avoid weight volatility. Here’s the breakdown:

  1. Larger inputs have more impact on the output
    The perceptron’s output hinges on the weighted sum $\sum w_j x_j + b$. If $x_i$ is large, even a small tweak to $w_i$ can shift the weighted sum (and thus the output) significantly. The update rule leverages this: when the perceptron makes a mistake ($d_i \neq o_i$), features that contributed most to that error (i.e., those with larger $|x_i|$) get adjusted more aggressively to fix the error fast.

  2. Updates are error-directed, not random
    Your concern about "small $w_i$ fluctuations causing big output changes" applies to uncontrolled weight shifts—but the perceptron update always moves in the direction of reducing the error $(d_i - o_i)$. For example:

    • If the perceptron underpredicts ($d_i=1$, $o_i=0$) and $x_i$ is a large positive value, increasing $w_i$ (via a positive $\Delta w_i$) will sharply boost the weighted sum, pushing the output closer to the expected value.
    • If $x_i$ were small, a smaller update makes sense because that feature wasn’t a major driver of the incorrect prediction.
  3. The training rate $\eta$ acts as a global safety knob
    If huge $x_i$ values feel like they’re causing overly drastic updates, the training rate $\eta$ is meant to scale all adjustments. You can reduce $\eta$ to make even large-$x_i$ updates more gradual, without breaking the core logic of prioritizing changes to impactful features.

A Quick Example to Illustrate

Suppose we have two input features: $x_1 = 10$, $x_2 = 1$. The perceptron’s error is $(d_i - o_i) = 1$, and $\eta = 0.1$.

  • $\Delta w_1 = 0.1 * 1 * 10 = 1$
  • $\Delta w_2 = 0.1 * 1 * 1 = 0.1$

After updating, the weighted sum increases by $101 + 10.1 = 10.1$. Since $x_1$ was the dominant feature in the original sum, this large adjustment fixes the error far faster than a tiny tweak to $w_1$ would.

In short: the update rule treats larger input magnitudes as a signal that the corresponding weight has more "responsibility" for the current error, so it adjusts that weight more to quickly get the prediction back on track.

内容的提问来源于stack exchange,提问作者demalegabi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:14:09