You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言手动实现感知器不收敛问题排查求助

Hey there! Let's figure out why your hand-built perceptron in R is stuck at ~70% accuracy and won't converge—especially since logistic regression is crushing it with 99%+ accuracy. That tells me your data is definitely linearly separable, so the issue must be in the implementation details. Let's break this down step by step.

1. First: Did You Include a Bias Term?

This is the #1 mistake I see with manual perceptron implementations. Logistic regression automatically includes an intercept (bias) term, but if you forgot to add one to your perceptron, your decision boundary can only pass through the origin. That’s a huge problem if your data’s decision boundary is shifted away from (0,0).

Fix this by adding a constant input of 1 to every sample, paired with a dedicated bias weight. This lets your model shift the decision boundary to fit the data properly.

2. Verify Your Error & Gradient Calculation

For a sigmoid-activated perceptron doing binary classification, cross-entropy loss is way more stable than mean squared error (MSE) for convergence. Even more importantly, you need to compute the gradient correctly for weight updates:

  • For a single sample, your predicted output is y_hat = sigmoid(w1*x1 + w2*x2 + b)
  • The cross-entropy loss gradient for each weight simplifies to (y_hat - y) * x (no need to multiply by the sigmoid derivative—cross-entropy cancels that out automatically!)
  • If you were using MSE, you’d need to multiply by y_hat*(1-y_hat) (the sigmoid derivative), but skipping this or messing up the sign will cause your weights to update in the wrong direction, leading to that back-and-forth fluctuation you’re seeing.

3. Check Your Weight Update Direction

Weight updates should always move in the direction that reduces loss: w = w - learning_rate * gradient. If you accidentally used + instead of -, you’re actively increasing loss each iteration, which explains the non-convergence.

4. Tweak Your Learning Rate (and Maybe Use Dynamic Decay)

A fixed learning rate that’s too big will make weights bounce around the optimal value; too small will make convergence crawl. Try starting with 0.01 or 0.001 (instead of 0.1 or higher). For even better stability, add dynamic decay—e.g., multiply the learning rate by 0.9 every 100 epochs to slow updates as you get closer to the optimal weights.


Corrected Perceptron Implementation Example

Here’s a fixed version of your code that addresses all these issues:

# Generate linearly separable synthetic data (same as your setup)
set.seed(123)
n <- 1000
x1 <- rnorm(n)
x2 <- rnorm(n)
y <- as.numeric(x1 + x2 > 0)  # Decision boundary: x1 + x2 > 0
data <- data.frame(x1, x2, y)

# Hand-coded perceptron with bias, cross-entropy gradient, and online updates
train_perceptron <- function(data, lr = 0.01, epochs = 1000) {
  # Initialize weights: w1, w2, bias (b)
  weights <- runif(3, -0.1, 0.1)
  w1 <- weights[1]
  w2 <- weights[2]
  b <- weights[3]
  
  for (epoch in 1:epochs) {
    # Online learning: update weights per sample
    for (i in 1:nrow(data)) {
      x1_i <- data$x1[i]
      x2_i <- data$x2[i]
      y_i <- data$y[i]
      
      # Compute weighted sum and sigmoid output
      z <- w1*x1_i + w2*x2_i + b
      y_hat <- 1 / (1 + exp(-z))
      
      # Calculate gradient (cross-entropy loss derivative)
      error <- y_hat - y_i
      dw1 <- error * x1_i
      dw2 <- error * x2_i
      db <- error
      
      # Update weights (move opposite to gradient)
      w1 <- w1 - lr * dw1
      w2 <- w2 - lr * dw2
      b <- b - lr * db
    }
    
    # Print progress every 100 epochs
    if (epoch %% 100 == 0) {
      z_all <- w1*data$x1 + w2*data$x2 + b
      y_hat_all <- ifelse(1/(1 + exp(-z_all)) > 0.5, 1, 0)
      acc <- mean(y_hat_all == data$y)
      cat(sprintf("Epoch %d | Training Accuracy: %.4f\n", epoch, acc))
    }
  }
  
  return(list(w1 = w1, w2 = w2, b = b))
}

# Train the model
trained_model <- train_perceptron(data, lr = 0.01, epochs = 1000)

# Evaluate final accuracy
z_test <- trained_model$w1*data$x1 + trained_model$w2*data$x2 + trained_model$b
y_hat_test <- ifelse(1/(1 + exp(-z_test)) > 0.5, 1, 0)
cat(sprintf("Final Training Accuracy: %.4f\n", mean(y_hat_test == data$y)))

Extra Debugging Tips

  • Plot loss/accuracy over epochs: Track the average loss or accuracy each epoch to see if it’s trending downward (this confirms convergence).
  • Standardize your inputs: Scaling x1 and x2 to have mean 0 and variance 1 can make learning more stable, especially if your features are on different scales.
  • Try batch gradient descent: Instead of updating weights per sample, compute the average gradient across all samples each epoch. This reduces noise in weight updates and can help with convergence.

内容的提问来源于stack exchange,提问作者Ju Ko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 11:12:32