如何用数学方法求解感知机的weight1、weight2及偏置?通用方法解析
Alright, let's break down how to calculate a perceptron's weight1, weight2, and bias—first with a concrete, small-scale example, then the general method that works for any problem.
First: The Perceptron Fundamentals
A perceptron's output follows this core rule:
y = sign(w₁x₁ + w₂x₂ + b)
Where:
sign(z)is the sign function: returns 1 if z > 0, -1 (or 0, depending on convention) if z ≤ 0x₁, x₂are your input featuresw₁, w₂are the weights we need to solve forbis the bias term (shifts the decision boundary left/right)
Case 1: Solving Small, Linearly Separable Problems (Direct Inequality Method)
For simple, linearly separable tasks (like logic gates: AND, OR, NOT), you can set up a system of inequalities based on your training data and find feasible parameter values.
Let’s use the AND gate as an example:
| x₁ | x₂ | y_true |
|---|---|---|
| 0 | 0 | -1 |
| 0 | 1 | -1 |
| 1 | 0 | -1 |
| 1 | 1 | 1 |
Translate each row into an inequality using the perceptron rule:
- For (0,0):
w₁*0 + w₂*0 + b ≤ 0→b ≤ 0 - For (0,1):
w₁*0 + w₂*1 + b ≤ 0→w₂ + b ≤ 0 - For (1,0):
w₁*1 + w₂*0 + b ≤ 0→w₁ + b ≤ 0 - For (1,1):
w₁*1 + w₂*1 + b > 0→w₁ + w₂ + b > 0
Now find values that satisfy all four. A common valid solution here is w₁=1, w₂=1, b=-1.5:
- Check (0,0): 0+0-1.5 = -1.5 ≤0 → correct
- Check (0,1):0+1-1.5=-0.5 ≤0 → correct
- Check (1,0):1+0-1.5=-0.5 ≤0 → correct
- Check (1,1):1+1-1.5=0.5>0 → correct
This method works great for tiny problems, but it’s not scalable for larger datasets.
Case 2: General Method for Any Problem (Perceptron Learning Algorithm)
For any perceptron problem—small or large, as long as the data is linearly separable—you’ll use the Perceptron Learning Algorithm, a tailored stochastic gradient descent approach for binary classification. Here’s the step-by-step process:
Initialize Parameters: Start with
w₁=0, w₂=0, b=0(or small random values to avoid symmetry issues). Pick a learning rateη(typically between 0.1 and 1—this controls how big each parameter update step is).Iterate Over Training Data: For each sample
(x₁, x₂, y_true):- Calculate the predicted output:
y_pred = sign(w₁x₁ + w₂x₂ + b) - If the prediction is wrong (
y_pred ≠ y_true), update the parameters:w₁ = w₁ + η * y_true * x₁ w₂ = w₂ + η * y_true * x₂ b = b + η * y_true - Why this works? When you misclassify,
y_true*(w₁x₁ + w₂x₂ + b) ≤ 0. Addingη*y_true*x₁tow₁pushes the decision boundary in the direction that will fix the misclassification next time.
- Calculate the predicted output:
Repeat Until Convergence: Keep looping through the data until all samples are classified correctly. For linearly separable data, this algorithm is guaranteed to converge (thanks to the Perceptron Convergence Theorem).
Handling Non-Linearly Separable Data
If your data isn’t linearly separable, the algorithm will never fully converge. In this case, use the Pocket Algorithm:
- Track the set of parameters (
w₁, w₂, b) that has the lowest number of misclassifications so far - Stop after a fixed number of iterations, then use the "pocket" parameters as your best solution
内容的提问来源于stack exchange,提问作者Rahul Vansh

