如何从多元正态分布生成训练实例并实现Python逻辑回归感知机
Hey there! Let's tackle your two questions one by one, with hands-on Python code you can test immediately.
In Python, the simplest way to generate samples from a multivariate normal distribution is using numpy.random.multivariate_normal. This function takes three core arguments:
mean: Your mean vector (likeμ1 = [1,0]orμ2 = [0,1.5])cov: The covariance matrix (yourΣ1orΣ2)size: The number of samples you want to generate per class
For example, to generate 1500 samples for your first class:
import numpy as np mu1 = [1, 0] sigma1 = [[1, 0.75], [0.75, 1]] class0_samples = np.random.multivariate_normal(mu1, sigma1, 1500)
You asked for a perceptron for logistic regression-style classification, using 3000 training/test samples split evenly between two multivariate normal classes. Let's build this end-to-end.
Step 1: Import Dependencies
We'll use NumPy for data generation and calculations, plus Matplotlib (optional) to visualize our data and model.
import numpy as np import matplotlib.pyplot as plt # Optional, for visualization
Step 2: Generate Training & Test Data
First, let's write a helper function to generate and shuffle our data (shuffling ensures our model doesn't learn order-based patterns):
def generate_class_data(mu1, sigma1, mu2, sigma2, samples_per_class): # Generate samples for class 0 (label 0) class0 = np.random.multivariate_normal(mu1, sigma1, samples_per_class) labels0 = np.zeros(samples_per_class) # Generate samples for class 1 (label 1) class1 = np.random.multivariate_normal(mu2, sigma2, samples_per_class) labels1 = np.ones(samples_per_class) # Combine and shuffle data all_samples = np.vstack((class0, class1)) all_labels = np.hstack((labels0, labels1)) shuffle_idx = np.random.permutation(len(all_samples)) return all_samples[shuffle_idx], all_labels[shuffle_idx] # Use your specified parameters mu1 = [1, 0] sigma1 = [[1, 0.75], [0.75, 1]] mu2 = [0, 1.5] sigma2 = [[1, 0.75], [0.75, 1]] # Generate 3000 training samples (1500 per class) and 3000 test samples X_train, y_train = generate_class_data(mu1, sigma1, mu2, sigma2, 1500) X_test, y_test = generate_class_data(mu1, sigma1, mu2, sigma2, 1500)
Step 3: Build the Perceptron Class
Perceptrons work best with labels encoded as -1 and 1 (instead of 0 and 1), so we'll adjust our labels during training. Here's a complete perceptron implementation with fit/predict methods:
class PerceptronClassifier: def __init__(self, learning_rate=0.01, num_epochs=1000): self.lr = learning_rate self.epochs = num_epochs self.weights = None self.bias = None def fit(self, X, y): n_samples, n_features = X.shape # Initialize weights to 0, bias to 0 self.weights = np.zeros(n_features) self.bias = 0 # Convert 0/1 labels to -1/1 for perceptron update rule y_encoded = np.where(y == 0, -1, 1) # Training loop for _ in range(self.epochs): for idx, sample in enumerate(X): # Compute linear combination of features and weights linear_out = np.dot(sample, self.weights) + self.bias # Predict label pred = np.sign(linear_out) # Update weights/bias if prediction is wrong if y_encoded[idx] != pred: self.weights += self.lr * y_encoded[idx] * sample self.bias += self.lr * y_encoded[idx] def predict(self, X): # Compute linear outputs linear_out = np.dot(X, self.weights) + self.bias # Convert back to 0/1 labels return np.where(np.sign(linear_out) == -1, 0, 1)
Step 4: Train & Evaluate the Model
Now let's train our perceptron and check its accuracy on the test data:
# Initialize and train the model perceptron = PerceptronClassifier(learning_rate=0.01, num_epochs=1000) perceptron.fit(X_train, y_train) # Make predictions on test data test_predictions = perceptron.predict(X_test) # Calculate accuracy accuracy = np.mean(test_predictions == y_test) print(f"Test Set Accuracy: {accuracy:.2f}")
Optional: Visualize the Data & Decision Boundary
To see how well our model separates the two classes, we can plot the data and the perceptron's decision boundary:
# Plot training data plt.scatter(X_train[:, 0], X_train[:, 1], c=y_train, cmap='viridis', alpha=0.6) plt.xlabel('Feature 1') plt.ylabel('Feature 2') plt.title('Training Data & Perceptron Decision Boundary') # Calculate decision boundary line x_range = np.linspace(-3, 4, 100) y_boundary = -(perceptron.weights[0] * x_range + perceptron.bias) / perceptron.weights[1] plt.plot(x_range, y_boundary, 'r--', label='Decision Boundary') plt.legend() plt.show()
内容的提问来源于stack exchange,提问作者user9036613

