分类器训练耗时过长求助:大尺寸音乐特征数据集优化方案
Hey there! Let's break down how to tackle this large feature array issue for your music genre classification task—you're on the right track, but we can optimize a few key areas to get training moving faster.
First off, I notice you're treating every MFCC frame as an individual sample, which is why your feature array blew up to 1.29M rows. That's not ideal for genre classification—we need to treat each song as a single sample, not each time frame.
Instead of stacking every frame's MFCC values, calculate statistical features across all frames of a song (like mean, std, min/max, percentiles). This compresses each song into a single feature vector, cutting your sample count from 1.29M to 1000 instantly.
Here's how to adjust your feature extraction code:
import numpy as np import librosa from sklearn.preprocessing import StandardScaler def extract_song_features(file_path, scaler=None, fit_scaler=False): y, sr = librosa.load(file_path) mfcc = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=20).T # Shape: (frames, 20) # Compute statistical features for each MFCC coefficient song_features = [] for mfcc_coeff in mfcc.T: song_features.extend([ np.mean(mfcc_coeff), np.std(mfcc_coeff), np.max(mfcc_coeff), np.min(mfcc_coeff), np.median(mfcc_coeff), np.percentile(mfcc_coeff, 25), np.percentile(mfcc_coeff, 75) ]) song_features = np.array(song_features).reshape(1, -1) # Shape: (1, 140) (20 coeffs ×7 stats) # Handle scaling correctly (no data leakage!) if fit_scaler and scaler is not None: return scaler.fit_transform(song_features) elif scaler is not None: return scaler.transform(song_features) return song_features # Example usage for your dataset scaler = StandardScaler() all_features = [] all_labels = [] dataset_files = [("path/to/song1.wav", 0), ("path/to/song2.wav", 1), ...] # Your (file, label) pairs # Fit scaler ONLY on training data (split your dataset first!) fit_scaler = True for file_path, label in dataset_files: if fit_scaler: feat = extract_song_features(file_path, scaler, fit_scaler=True) fit_scaler = False # Stop fitting after training data else: feat = extract_song_features(file_path, scaler) all_features.append(feat) all_labels.append(label) # Final feature array shape: (1000, 140) — way more manageable! X = np.vstack(all_features) y = np.array(all_labels)
TSNE is great for visualization, but it’s computationally expensive and not designed for preprocessing training data. Stick to faster, more practical methods:
- PCA: Fast, linear, and retains most signal by reducing correlated features. Use it to keep 95% of variance (adjust as needed):
from sklearn.decomposition import PCA pca = PCA(n_components=0.95) # Retain 95% of variance X_pca = pca.fit_transform(X) - TruncatedSVD: Similar to PCA but faster on large datasets, especially if your features are sparse.
- Feature Selection: Use
SelectKBestwith ANOVA F-values to keep only the most predictive features, cutting dimensionality without losing critical signal.
- Start with Lightweight Models: For structured audio features (like our song-level stats), linear models like Logistic Regression or Linear SVM train way faster than deep learning models and often perform well:
from sklearn.linear_model import LogisticRegression from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split(X_pca, y, test_size=0.2) model = LogisticRegression(max_iter=1000, n_jobs=-1) # n_jobs uses all CPU cores model.fit(X_train, y_train) - Mini-Batch Training (for Deep Learning): If you want to use a neural network, use mini-batches instead of full-batch training. Keras/TensorFlow does this automatically with
batch_size:from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense from tensorflow.keras.callbacks import EarlyStopping model = Sequential([ Dense(64, activation='relu', input_shape=(X_pca.shape[1],)), Dense(32, activation='relu'), Dense(10, activation='softmax') ]) model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy']) # Add early stopping to save time and prevent overfitting early_stopping = EarlyStopping(patience=3, restore_best_weights=True) model.fit(X_train, y_train, batch_size=64, epochs=20, validation_split=0.1, callbacks=[early_stopping]) - GPU Acceleration: If you’re using deep learning, make sure your framework (TensorFlow/Keras) is using a GPU—this will drastically speed up training.
- Avoid Data Leakage: Never fit your scaler on the entire dataset—fit only on training data, then transform both train and test sets. Your original code used
fit_transformper song, which is incorrect. - Parallel Feature Extraction: Use
joblibto extract features from multiple songs at once, saving preprocessing time:from joblib import Parallel, delayed def process_song(file_path, label): return extract_song_features(file_path), label # Process 8 songs in parallel (adjust n_jobs based on your CPU cores) results = Parallel(n_jobs=8)(delayed(process_song)(fp, lbl) for fp, lbl in dataset_files) all_features = [res[0] for res in results] all_labels = [res[1] for res in results]
The biggest win here is switching from frame-level samples to song-level statistical features—this cuts your dataset size by over 99% and makes training orders of magnitude faster. Pair that with the optimizations above, and you’ll be able to train your classifier in no time.
内容的提问来源于stack exchange,提问作者user193713

