基于Python的DBSCAN聚类:为鸢尾花数据集添加DBSCAN聚类列并导出
Got it, let's walk through exactly how to add that DBSCAN column to your dataset and export the updated file. Here's a step-by-step breakdown with code tailored to your specific needs:
Step 1: Import Required Libraries
We’ll need pandas for handling CSV data and DBSCAN from scikit-learn for clustering:
import pandas as pd from sklearn.cluster import DBSCAN
Step 2: Load Your Dataset
Use a raw string for your Windows file path (the r prefix) to avoid backslash-related errors:
# Load the existing CSV file file_path = r"C:\Users\05807\Desktop\New folder (5)\FinalX.csv" df = pd.read_csv(file_path)
Step 3: Prepare Numerical Features for DBSCAN
DBSCAN only works with numerical data, so we’ll select the four iris measurement columns as our clustering inputs:
# Pick numerical columns for clustering (exclude text/ID/existing cluster columns) numerical_features = df[["Sepal_length", "Sepal_width", "Petal_length", "Petal_width"]]
Step 4: Run DBSCAN and Format Cluster Labels
By default, DBSCAN returns integer labels (0, 1, 2, etc., plus -1 for noise points). We’ll convert these to the "Cluster X" format you want, and clearly mark noise:
# Initialize and fit DBSCAN (tweak eps/min_samples to get your desired 3 clusters) dbscan = DBSCAN(eps=0.5, min_samples=5) cluster_labels = dbscan.fit_predict(numerical_features) # Convert labels to human-readable format formatted_labels = [] for label in cluster_labels: if label == -1: formatted_labels.append("Noise") else: formatted_labels.append(f"Cluster {label + 1}") # Start counting from 1 instead of 0
Note: Adjust eps (maximum distance between points in a cluster) and min_samples (minimum points required to form a cluster) if you don’t get 3 distinct clusters initially—test small changes to find the best fit for your data.
Step 5: Insert the DBSCAN Column After "Hierarchical"
We’ll place the new column exactly where you want it:
# Find the position of the "Hierarchical" column hierarchical_col_index = df.columns.get_loc("Hierarchical") # Insert the DBSCAN column right after it df.insert(hierarchical_col_index + 1, "DBSCAN", formatted_labels)
Step 6: Export the Updated Dataset
Save the modified data to a new CSV file:
# Define output path (adjust this to your preferred location/filename) output_path = r"C:\Users\05807\Desktop\New folder (5)\FinalX_With_DBSCAN.csv" df.to_csv(output_path, index=False) print(f"Updated dataset saved successfully to: {output_path}")
Full Working Code
Putting it all together, here’s the complete script you can run:
import pandas as pd from sklearn.cluster import DBSCAN # Load original dataset file_path = r"C:\Users\05807\Desktop\New folder (5)\FinalX.csv" df = pd.read_csv(file_path) # Select numerical features for clustering numerical_features = df[["Sepal_length", "Sepal_width", "Petal_length", "Petal_width"]] # Run DBSCAN clustering dbscan = DBSCAN(eps=0.5, min_samples=5) cluster_labels = dbscan.fit_predict(numerical_features) # Format labels to "Cluster X" style formatted_labels = [] for label in cluster_labels: if label == -1: formatted_labels.append("Noise") else: formatted_labels.append(f"Cluster {label + 1}") # Insert DBSCAN column after Hierarchical hierarchical_col_index = df.columns.get_loc("Hierarchical") df.insert(hierarchical_col_index + 1, "DBSCAN", formatted_labels) # Export updated data to new CSV output_path = r"C:\Users\05807\Desktop\New folder (5)\FinalX_With_DBSCAN.csv" df.to_csv(output_path, index=False) print("Done! Your dataset now includes the DBSCAN cluster column.")
Just execute this script, and you’ll have your new CSV with the DBSCAN column positioned exactly after "Hierarchical".
内容的提问来源于stack exchange,提问作者Snyder Fox

