如何用Python读取CSV数据集并应用Gradient descent完成精度评估与可视化
Got it, let's walk through adapting your existing gradient descent code to work with your CSV dataset, plus add proper accuracy evaluation and fix up the visualization. I'll break this down step by step with a full modified code example.
1. Reading Your CSV Dataset
First, we'll replace the synthetic make_blobs data with your CSV. I'll use pandas (install it with pip install pandas if you haven't) because it's perfect for loading tabular data. Assumptions: Your CSV has:
- Features in the first N columns
- Binary labels (0/1) in the last column
- No missing values (we can add handling if needed, but let's start simple)
2. Adding Accuracy Evaluation
After training, we'll calculate how many predictions match the true labels to get an accuracy score—this is a basic but critical metric for classification tasks.
3. Fixing the Visualization
The original code had a small bug in calculating the decision boundary. We'll correct that to plot a clean separating line between classes.
Full Modified Code
import matplotlib.pyplot as plt import numpy as np import argparse import pandas as pd # Add pandas for CSV reading def sigmoid_activation(x): return 1.0 / (1 + np.exp(-x)) ap = argparse.ArgumentParser() ap.add_argument("-e", "--epochs", type=int, default=100, help="# of epochs") ap.add_argument("-a", "--alpha", type=float, default=0.01, help="learning rate") ap.add_argument("-d", "--dataset", required=True, help="Path to your CSV dataset") # Add CSV path argument args = vars(ap.parse_args()) # -------------------------- # Step 1: Load your CSV data # -------------------------- print("[INFO] loading dataset...") df = pd.read_csv(args["dataset"]) # Split features (all columns except last) and labels (last column) X = df.iloc[:, :-1].values y = df.iloc[:, -1].values # Add bias term (column of 1s) to X, same as original code X = np.c_[np.ones((X.shape[0])), X] # -------------------------- # Training (mostly original, kept intact) # -------------------------- print("[INFO] starting training...") W = np.random.uniform(size=(X.shape[1],)) lossHistory = [] for epoch in np.arange(0, args["epochs"]): preds = sigmoid_activation(X.dot(W)) error = preds - y loss = np.sum(error ** 2) lossHistory.append(loss) print("[INFO] epoch #{}, loss={:.7f}".format(epoch + 1, loss)) gradient = X.T.dot(error) / X.shape[0] W += -args["alpha"] * gradient # -------------------------- # Step 2: Calculate accuracy # -------------------------- print("\n[INFO] evaluating model...") # Get predicted labels (0 if activation < 0.5, else 1) predicted_labels = np.where(sigmoid_activation(X.dot(W)) < 0.5, 0, 1) # Calculate accuracy: percentage of correct predictions accuracy = np.mean(predicted_labels == y) * 100 print("[INFO] model accuracy: {:.2f}%".format(accuracy)) # Optional: Print a few sample predictions (like original code) print("\n[INFO] sample predictions:") sample_indices = np.random.choice(X.shape[0], 10) for i in sample_indices: activation = sigmoid_activation(X[i].dot(W)) label = 0 if activation < 0.5 else 1 print("activation={:.4f}; predicted_label={}, true_label={}".format(activation, label, y[i])) # -------------------------- # Step 3: Fixed Visualization # -------------------------- # Plot data points and decision boundary plt.figure(figsize=(8, 6)) # Plot features (assuming 2 features for visualization; adjust if you have more) plt.scatter(X[:, 1], X[:, 2], marker="o", c=y, edgecolor="k") # Calculate decision boundary: x2 = (-W[0] - W[1]*x1)/W[2] x_min, x_max = X[:, 1].min() - 1, X[:, 1].max() + 1 x1 = np.linspace(x_min, x_max, 100) x2 = (-W[0] - W[1] * x1) / W[2] plt.plot(x1, x2, "r-", label="Decision Boundary") plt.title("Dataset with Decision Boundary") plt.xlabel("Feature 1") plt.ylabel("Feature 2") plt.legend() # Plot training loss fig = plt.figure(figsize=(8, 6)) plt.plot(np.arange(0, args["epochs"]), lossHistory) fig.suptitle("Training Loss") plt.xlabel("Epoch #") plt.ylabel("Loss") plt.show()
How to Use This Code
- Save the code as
gradient_descent_csv.py - Run it with your CSV path:
python gradient_descent_csv.py --dataset path/to/your/data.csv --epochs 200 --alpha 0.05
Notes
- If your dataset has more than 2 features, the scatter plot won't work (you can't plot 3+ features easily). In that case, you can skip the decision boundary plot or use dimensionality reduction (like PCA) first.
- If your labels are not 0/1, adjust the code to binarize them (e.g., map class A to 0, class B to 1) before training.
- If you have missing values in your CSV, add
df = df.dropna()after loading to remove rows with missing data, or use imputation.
内容的提问来源于stack exchange,提问作者Deema20

