如何在Python中绘制逻辑回归θ₁、θ₂的对数似然函数以判断凸性
Got it, let's walk through how to plot the negative log-likelihood function for your two-parameter logistic regression model and verify if it's convex. Here's a step-by-step guide with working Python code:
First, let's fix a small bug in your existing loss code—you forgot to include the (1 - y) * np.log(1 - sigma) term inside the np.sum() call. We need to sum the full loss for every sample, so the corrected formula is:
loss = -1 / size * np.sum(y * np.log(sigma) + (1 - y) * np.log(1 - sigma))
This ensures we're calculating the average negative log-likelihood across all data points correctly.
To plot the function, we need some sample input features x and binary labels y. Let's generate synthetic data with 2 features (since we're focusing on θ₁ and θ₂) and 100 samples:
import numpy as np import matplotlib.pyplot as plt from mpl_toolkits.mplot3d import Axes3D # Generate synthetic data np.random.seed(42) # For reproducibility n_samples = 100 x = np.random.randn(n_samples, 2) # 2 features, n_samples samples true_weights = [1.5, -0.8] # True θ₁ and θ₂ for generating labels z = np.dot(x, true_weights) y = np.random.binomial(1, 1 / (1 + np.exp(-z))) # Binary labels
We'll generate a grid of θ₁ and θ₂ values to evaluate the loss function across. Let's pick a range around the true weights to get a clear view:
# Define range for θ₁ and θ₂ theta1_range = np.linspace(-3, 3, 50) theta2_range = np.linspace(-3, 3, 50) theta1_grid, theta2_grid = np.meshgrid(theta1_range, theta2_range) # Initialize loss grid loss_grid = np.zeros_like(theta1_grid) n_samples = x.shape[0]
Loop through each (θ₁, θ₂) pair in the grid, compute the sigmoid and corresponding negative log-likelihood:
def sigmoid(z): return 1 / (1 + np.exp(-z)) for i in range(theta1_grid.shape[0]): for j in range(theta1_grid.shape[1]): weights = np.array([theta1_grid[i, j], theta2_grid[i, j]]) sigma = sigmoid(np.dot(x, weights)) # Avoid log(0) by adding a tiny epsilon epsilon = 1e-10 sigma = np.clip(sigma, epsilon, 1 - epsilon) loss = -1 / n_samples * np.sum(y * np.log(sigma) + (1 - y) * np.log(1 - sigma)) loss_grid[i, j] = loss
Note: We add a small epsilon and clip the sigmoid values to avoid np.log(0) errors, which would cause infinite loss values.
We can plot this as a 3D surface or a contour plot to visualize its shape:
3D Surface Plot
fig = plt.figure(figsize=(10, 7)) ax = fig.add_subplot(111, projection='3d') surf = ax.plot_surface(theta1_grid, theta2_grid, loss_grid, cmap='viridis') ax.set_xlabel('θ₁') ax.set_ylabel('θ₂') ax.set_zlabel('Negative Log-Likelihood') ax.set_title('Negative Log-Likelihood of Logistic Regression (θ₁ vs θ₂)') fig.colorbar(surf, shrink=0.5, aspect=5) plt.show()
Contour Plot (2D Top-Down View)
plt.figure(figsize=(8, 6)) contour = plt.contourf(theta1_grid, theta2_grid, loss_grid, cmap='viridis', levels=20) plt.xlabel('θ₁') plt.ylabel('θ₂') plt.title('Contour Plot of Negative Log-Likelihood') plt.colorbar(contour) plt.show()
When you plot the function, you'll see it forms a smooth, bowl-shaped surface (or convex contours in 2D). This confirms that the negative log-likelihood for logistic regression is indeed a convex function—it has no local minima, only a single global minimum, which is why gradient descent works reliably for training logistic regression models.
内容的提问来源于stack exchange,提问作者AfonsoSalgadoSousa

