如何为多文件夹手势CSV数据集构建训练测试文件并训练Python分类器
Hey Danny, let's walk through this step by step — I've worked with similar CSV-based gesture datasets before, so this process should feel straightforward once we break it down. Here's how to turn your folder-structured data into trainable features/labels, then train and test classifiers in Python:
Step 1: Load & Preprocess Your Dataset
First, we'll iterate through your folder structure to build two lists: one for feature vectors (each from a CSV file) and one for corresponding labels (the folder name each CSV belongs to). We'll also convert text labels to numeric values since most classifiers require this.
import os import pandas as pd from sklearn.preprocessing import LabelEncoder # Replace this with your actual dataset root path dataset_root = "./your_gesture_dataset" features = [] raw_labels = [] # Loop through each gesture folder for gesture_dir in os.listdir(dataset_root): dir_path = os.path.join(dataset_root, gesture_dir) if not os.path.isdir(dir_path): continue # Skip any non-folder files # Loop through all CSV files in the gesture folder for csv_file in os.listdir(dir_path): if not csv_file.endswith(".csv"): continue csv_path = os.path.join(dir_path, csv_file) # Read CSV and flatten into a 1D feature vector # Add header=None if your CSV has no column headers df = pd.read_csv(csv_path) feature_vector = df.values.flatten() # Turns 2D table into 1D array features.append(feature_vector) raw_labels.append(gesture_dir) # Use folder name as label # Convert text labels to numeric values label_encoder = LabelEncoder() numeric_labels = label_encoder.fit_transform(raw_labels)
Step 2: Split Into Training & Testing Sets
We'll split our data into training (for teaching the model) and testing (for evaluating performance) subsets using scikit-learn's built-in tool. This ensures we test on unseen data to get a realistic idea of model accuracy.
from sklearn.model_selection import train_test_split # Split 80% of data for training, 20% for testing # random_state ensures results are reproducible X_train, X_test, y_train, y_test = train_test_split( features, numeric_labels, test_size=0.2, random_state=42 )
Step 3: Train Different Classifiers
Now let's train a few common classifiers. You can experiment with all of these to see which works best for your gesture data:
Option 1: Support Vector Machine (SVM)
Great for high-dimensional data like gesture features:
from sklearn.svm import SVC svm_classifier = SVC(kernel='rbf') # RBF kernel handles non-linear patterns; try 'linear' too svm_classifier.fit(X_train, y_train)
Option 2: Random Forest
Robust to overfitting and easy to tune:
from sklearn.ensemble import RandomForestClassifier rf_classifier = RandomForestClassifier(n_estimators=100, random_state=42) rf_classifier.fit(X_train, y_train)
Option 3: K-Nearest Neighbors (KNN)
Simple and intuitive for pattern matching:
from sklearn.neighbors import KNeighborsClassifier knn_classifier = KNeighborsClassifier(n_neighbors=5) # Adjust n_neighbors based on your data knn_classifier.fit(X_train, y_train)
Step 4: Test & Evaluate Model Performance
Once trained, we'll test each model on the test set and measure how well it performs:
from sklearn.metrics import accuracy_score, classification_report # Example with SVM (repeat for other classifiers) y_pred = svm_classifier.predict(X_test) # Print overall accuracy print(f"SVM Accuracy: {accuracy_score(y_test, y_pred):.2f}") # Print detailed report (precision, recall, F1-score per gesture class) print("\nSVM Classification Report:") print(classification_report(y_test, y_pred, target_names=label_encoder.classes_)) # For Random Forest, just swap in the classifier: # y_pred_rf = rf_classifier.predict(X_test) # print(f"Random Forest Accuracy: {accuracy_score(y_test, y_pred_rf):.2f}")
Quick Tips for Better Results
- If your CSV has header rows, add
header=Nonetopd.read_csv()to avoid treating headers as features. - If feature dimensions are very high, use PCA to reduce noise:
from sklearn.decomposition import PCA pca = PCA(n_components=0.95) # Keep 95% of variance X_train_pca = pca.fit_transform(X_train) X_test_pca = pca.transform(X_test) # Train classifiers on X_train_pca instead - Use cross-validation for more reliable performance estimates:
from sklearn.model_selection import cross_val_score svm_scores = cross_val_score(svm_classifier, features, numeric_labels, cv=5) print(f"SVM Cross-Validation Scores: {svm_scores}")
内容的提问来源于stack exchange,提问作者Danny

