You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分类器训练耗时过长求助:大尺寸音乐特征数据集优化方案

Hey there! Let's break down how to tackle this large feature array issue for your music genre classification task—you're on the right track, but we can optimize a few key areas to get training moving faster.

1. Fix the Sample Definition (Critical!)

First off, I notice you're treating every MFCC frame as an individual sample, which is why your feature array blew up to 1.29M rows. That's not ideal for genre classification—we need to treat each song as a single sample, not each time frame.

Instead of stacking every frame's MFCC values, calculate statistical features across all frames of a song (like mean, std, min/max, percentiles). This compresses each song into a single feature vector, cutting your sample count from 1.29M to 1000 instantly.

Here's how to adjust your feature extraction code:

import numpy as np
import librosa
from sklearn.preprocessing import StandardScaler

def extract_song_features(file_path, scaler=None, fit_scaler=False):
    y, sr = librosa.load(file_path)
    mfcc = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=20).T  # Shape: (frames, 20)
    
    # Compute statistical features for each MFCC coefficient
    song_features = []
    for mfcc_coeff in mfcc.T:
        song_features.extend([
            np.mean(mfcc_coeff),
            np.std(mfcc_coeff),
            np.max(mfcc_coeff),
            np.min(mfcc_coeff),
            np.median(mfcc_coeff),
            np.percentile(mfcc_coeff, 25),
            np.percentile(mfcc_coeff, 75)
        ])
    song_features = np.array(song_features).reshape(1, -1)  # Shape: (1, 140) (20 coeffs ×7 stats)
    
    # Handle scaling correctly (no data leakage!)
    if fit_scaler and scaler is not None:
        return scaler.fit_transform(song_features)
    elif scaler is not None:
        return scaler.transform(song_features)
    return song_features

# Example usage for your dataset
scaler = StandardScaler()
all_features = []
all_labels = []
dataset_files = [("path/to/song1.wav", 0), ("path/to/song2.wav", 1), ...]  # Your (file, label) pairs

# Fit scaler ONLY on training data (split your dataset first!)
fit_scaler = True
for file_path, label in dataset_files:
    if fit_scaler:
        feat = extract_song_features(file_path, scaler, fit_scaler=True)
        fit_scaler = False  # Stop fitting after training data
    else:
        feat = extract_song_features(file_path, scaler)
    all_features.append(feat)
    all_labels.append(label)

# Final feature array shape: (1000, 140) — way more manageable!
X = np.vstack(all_features)
y = np.array(all_labels)
2. Choose the Right Dimensionality Reduction Methods

TSNE is great for visualization, but it’s computationally expensive and not designed for preprocessing training data. Stick to faster, more practical methods:

  • PCA: Fast, linear, and retains most signal by reducing correlated features. Use it to keep 95% of variance (adjust as needed):
    from sklearn.decomposition import PCA
    pca = PCA(n_components=0.95)  # Retain 95% of variance
    X_pca = pca.fit_transform(X)
    
  • TruncatedSVD: Similar to PCA but faster on large datasets, especially if your features are sparse.
  • Feature Selection: Use SelectKBest with ANOVA F-values to keep only the most predictive features, cutting dimensionality without losing critical signal.
3. Optimize Model Training
  • Start with Lightweight Models: For structured audio features (like our song-level stats), linear models like Logistic Regression or Linear SVM train way faster than deep learning models and often perform well:
    from sklearn.linear_model import LogisticRegression
    from sklearn.model_selection import train_test_split
    
    X_train, X_test, y_train, y_test = train_test_split(X_pca, y, test_size=0.2)
    model = LogisticRegression(max_iter=1000, n_jobs=-1)  # n_jobs uses all CPU cores
    model.fit(X_train, y_train)
    
  • Mini-Batch Training (for Deep Learning): If you want to use a neural network, use mini-batches instead of full-batch training. Keras/TensorFlow does this automatically with batch_size:
    from tensorflow.keras.models import Sequential
    from tensorflow.keras.layers import Dense
    from tensorflow.keras.callbacks import EarlyStopping
    
    model = Sequential([
        Dense(64, activation='relu', input_shape=(X_pca.shape[1],)),
        Dense(32, activation='relu'),
        Dense(10, activation='softmax')
    ])
    model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
    
    # Add early stopping to save time and prevent overfitting
    early_stopping = EarlyStopping(patience=3, restore_best_weights=True)
    model.fit(X_train, y_train, batch_size=64, epochs=20, validation_split=0.1, callbacks=[early_stopping])
    
  • GPU Acceleration: If you’re using deep learning, make sure your framework (TensorFlow/Keras) is using a GPU—this will drastically speed up training.
4. Code & Data Pipeline Optimizations
  • Avoid Data Leakage: Never fit your scaler on the entire dataset—fit only on training data, then transform both train and test sets. Your original code used fit_transform per song, which is incorrect.
  • Parallel Feature Extraction: Use joblib to extract features from multiple songs at once, saving preprocessing time:
    from joblib import Parallel, delayed
    
    def process_song(file_path, label):
        return extract_song_features(file_path), label
    
    # Process 8 songs in parallel (adjust n_jobs based on your CPU cores)
    results = Parallel(n_jobs=8)(delayed(process_song)(fp, lbl) for fp, lbl in dataset_files)
    all_features = [res[0] for res in results]
    all_labels = [res[1] for res in results]
    

The biggest win here is switching from frame-level samples to song-level statistical features—this cuts your dataset size by over 99% and makes training orders of magnitude faster. Pair that with the optimizations above, and you’ll be able to train your classifier in no time.

内容的提问来源于stack exchange,提问作者user193713

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:23:09