Logistic回归代码问题排查及数据与决策边界绘制方法咨询
Hi there! No worries at all about your English—we're here to help you work through this logistic regression issue. Let's start by breaking down the problems in your code and then cover how to plot your data and decision boundary.
First, Let's Fix Your Logistic Regression Code
Looking at your snippet, there are a few key issues that are likely causing poor performance:
- Imports inside the class constructor: It's not a good practice to import
numpyinside__init__. Move imports to the top of your script instead—this keeps code clean and avoids redundant imports. - Incomplete sigmoid function: Your
sigFuncmethod is cut off! The sigmoid (core to logistic regression) needs to return1 / (1 + np.exp(-z)). - Missing training logic: There's no code to train the model (update weights/bias via gradient descent)—without this, your model can't learn from data.
- Missing prediction methods: You need functions to output either class labels or predicted probabilities after training.
Here's a complete, corrected version of your class:
import numpy as np class LogisticRegression(object): def __init__(self, eta=0.01, n_iter=1000): self.eta = eta # Learning rate self.n_iter = n_iter # Number of training loops self.weights = None self.bias = None def sigFunc(self, z): # Sigmoid function maps values to [0,1] range return 1 / (1 + np.exp(-z)) def fit(self, X, y): # Initialize weights and bias to 0 n_samples, n_features = X.shape self.weights = np.zeros(n_features) self.bias = 0 # Gradient descent training loop for _ in range(self.n_iter): # Calculate linear combination of inputs and weights linear_model = np.dot(X, self.weights) + self.bias # Convert to probabilities with sigmoid y_pred_proba = self.sigFunc(linear_model) # Compute gradients for weights and bias dw = (1 / n_samples) * np.dot(X.T, (y_pred_proba - y)) db = (1 / n_samples) * np.sum(y_pred_proba - y) # Update weights and bias self.weights -= self.eta * dw self.bias -= self.eta * db def predict_proba(self, X): # Return predicted probabilities for each class linear_model = np.dot(X, self.weights) + self.bias return self.sigFunc(linear_model) def predict(self, X): # Return hard class labels (0 or 1) proba = self.predict_proba(X) return np.where(proba >= 0.5, 1, 0)
How to Plot Data & Decision Boundary
We'll use matplotlib for visualization. Here's a step-by-step example using a sample dataset (replace it with your own data):
Step 1: Prepare Data & Train the Model
# Example dataset (swap with your actual data) from sklearn.datasets import make_classification X, y = make_classification(n_samples=200, n_features=2, n_informative=2, n_redundant=0, random_state=42) # Initialize and train the model model = LogisticRegression(eta=0.01, n_iter=1000) model.fit(X, y)
Step 2: Plot the Raw Data Points
import matplotlib.pyplot as plt # Plot each class with distinct colors plt.scatter(X[y == 0][:, 0], X[y == 0][:, 1], label='Class 0', c='navy') plt.scatter(X[y == 1][:, 0], X[y == 1][:, 1], label='Class 1', c='orange') plt.xlabel('Feature 1') plt.ylabel('Feature 2') plt.legend()
Step 3: Add the Decision Boundary
The decision boundary is where the model predicts a 50% probability of either class (sigmoid(0) = 0.5), which translates to the line w1*x1 + w2*x2 + b = 0.
# Create a grid of points to map the boundary x_min, x_max = X[:, 0].min() - 1, X[:, 0].max() + 1 y_min, y_max = X[:, 1].min() - 1, X[:, 1].max() + 1 xx, yy = np.meshgrid(np.arange(x_min, x_max, 0.01), np.arange(y_min, y_max, 0.01)) # Predict class for every point in the grid Z = model.predict(np.c_[xx.ravel(), yy.ravel()]) Z = Z.reshape(xx.shape) # Plot the boundary as a filled contour plt.contourf(xx, yy, Z, alpha=0.2, cmap=plt.cm.Paired) plt.show()
Quick Tips to Boost Model Performance
- Normalize your data: Logistic regression performs better when features are scaled (e.g., subtract the mean and divide by standard deviation for each feature).
- Tune hyperparameters: Adjust
eta(learning rate) andn_iter(training loops)—if the model isn't converging, try increasingn_iteror using a smallereta. - Check class balance: If one class has way more samples than the other, consider using class weights or resampling to fix imbalance.
If you run into specific error messages or have details about your dataset, feel free to share more!
内容的提问来源于stack exchange,提问作者user9641842

