分类器训练耗时过长求助:大尺寸MFCC特征数据优化方案
Hey there, let's work through this music genre classification bottleneck you're hitting. That massive (1.2M, 20) feature array is definitely a problem—t-SNE and cutting samples aren't the right fixes here, so let's dive into practical, efficient solutions that actually work.
First, a quick critical note: Your current scaler usage has a data leak—you’re fitting the scaler to each individual song’s MFCCs instead of the entire training dataset. That means each song is normalized using its own stats, which won’t generalize to test data. We’ll fix that in the examples below.
t-SNE is awesome for making pretty visualizations of your data, but it’s terrible for training pipeline preprocessing. It doesn’t preserve global data structure, and it’s computationally expensive (which is why you didn’t see speed gains). Stick to dimensionality reduction methods designed for training, like PCA or feature selection.
The root of your problem is treating every single MFCC frame as a separate training sample—music genre is a song-level trait, not a frame-level trait. Instead, collapse each song’s (1293,20) MFCC array into a fixed-length feature vector using statistical aggregates. This reduces your dataset from 1.2M samples to 1000 samples instantly.
Here’s how to implement this correctly:
import numpy as np import librosa from sklearn.preprocessing import StandardScaler from sklearn.model_selection import train_test_split from sklearn.ensemble import RandomForestClassifier # Define a function to extract aggregated song-level MFCC features def extract_song_mfcc_features(file_path, n_mfcc=20): y, sr = librosa.load(file_path) # Extract MFCCs (shape: (n_mfcc, n_frames)) mfcc = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=n_mfcc) # Calculate key statistics for each MFCC band stats_per_band = [] for mfcc_band in mfcc: stats_per_band.extend([ np.mean(mfcc_band), np.var(mfcc_band), np.max(mfcc_band), np.min(mfcc_band), np.median(mfcc_band), np.percentile(mfcc_band, 75) ]) return np.array(stats_per_band) # Assume you have lists of your song paths and corresponding genre labels (0-9) file_paths = [...] # Your 1000 song file paths labels = [...] # Corresponding genre labels # Extract features for all songs all_features = [extract_song_mfcc_features(path) for path in file_paths] all_features = np.array(all_features) # Shape: (1000, 120) (20 bands × 6 stats) # Standardize correctly (fit on all training data first) scaler = StandardScaler() scaled_features = scaler.fit_transform(all_features) # Split into train/test sets X_train, X_test, y_train, y_test = train_test_split( scaled_features, labels, test_size=0.2, random_state=42 ) # Train a classifier (Random Forest example; swap for your neural network if needed) clf = RandomForestClassifier(n_estimators=100, random_state=42) clf.fit(X_train, y_train) print(f"Test Accuracy: {clf.score(X_test, y_test):.4f}")
This approach works because genre is defined by consistent patterns across an entire song—stats like mean MFCC capture the overall timbre, while variance captures rhythmic dynamics.
If you need to use frame-level data (e.g., for sequence models like LSTMs/CNNs), use PCA to reduce the feature dimension without losing critical information:
from sklearn.decomposition import PCA # Assume your original frame-level features are (1293000, 20) pca = PCA(n_components=10) # Reduce to 10 dimensions X_pca = pca.fit_transform(original_frame_features) print(f"Total Explained Variance: {np.sum(pca.explained_variance_ratio_):.4f}") # Now use X_pca for training—half the feature size, way faster to process
If frame-level data is non-negotiable and still too big to fit in memory, use a data generator to load and process batches of songs on-the-fly. This avoids loading the entire dataset at once:
from tensorflow.keras.utils import Sequence class MusicFrameGenerator(Sequence): def __init__(self, file_paths, labels, batch_size=32, n_mfcc=20, scaler=None): self.file_paths = file_paths self.labels = labels self.batch_size = batch_size self.n_mfcc = n_mfcc self.scaler = scaler self.indexes = np.arange(len(file_paths)) def __len__(self): return int(np.ceil(len(self.file_paths) / self.batch_size)) def __getitem__(self, idx): batch_idx = self.indexes[idx*self.batch_size : (idx+1)*self.batch_size] batch_paths = [self.file_paths[i] for i in batch_idx] batch_labels = [self.labels[i] for i in batch_idx] # Load and process each song's MFCC frames batch_features = [] for path in batch_paths: y, sr = librosa.load(path) mfcc = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=self.n_mfcc).T if self.scaler: mfcc = self.scaler.transform(mfcc) # Pad/truncate frames to a fixed length (e.g., 1000 frames) mfcc_padded = np.zeros((1000, self.n_mfcc)) mfcc_padded[:min(len(mfcc), 1000)] = mfcc[:min(len(mfcc), 1000)] batch_features.append(mfcc_padded) return np.array(batch_features), np.array(batch_labels) # Initialize generators (fit scaler on training data first!) train_generator = MusicFrameGenerator(train_paths, train_labels, batch_size=32, scaler=scaler) val_generator = MusicFrameGenerator(val_paths, val_labels, batch_size=32, scaler=scaler) # Train your sequence model with the generator model.fit(train_generator, validation_data=val_generator, epochs=20)
- Use lightweight models: If using neural networks, avoid overly deep dense layers. For sequence data, use CNNs (they’re faster than LSTMs at extracting local temporal features) or small transformers.
- Faster optimizers: Swap SGD for AdamW—it converges faster and avoids overfitting better.
- Mixed precision training: Enable mixed precision in TensorFlow/Keras to reduce memory usage and speed up training:
import tensorflow as tf tf.keras.mixed_precision.set_global_policy('mixed_float16') - Early stopping: Add
EarlyStoppingto stop training once validation loss stops improving, saving unnecessary computation.
内容的提问来源于stack exchange,提问作者user45581

