You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将MP3转频率振幅数组并构建音乐评分神经网络?

Hey there! Let's break down your questions step by step—I’ve worked on similar audio ML projects before, so I’ll share practical, actionable solutions.


Q1: How to convert an MP3 file to frequency and amplitude arrays using Python?

For audio processing in Python, librosa is the go-to library—it’s built specifically for music and audio analysis, making it perfect for this task. Here’s a straightforward implementation:

First, install the library:

pip install librosa

Then, use this code to extract amplitude (time-domain) and frequency-domain amplitude data:

import librosa
import numpy as np

# Load the MP3 file; sr=None preserves the original sample rate of the audio
audio_waveform, sample_rate = librosa.load("your_audio_file.mp3", sr=None)

# `audio_waveform` is your amplitude array (time-domain)
# Each value represents the audio signal's amplitude at a specific time step
amplitude_array = audio_waveform

# Convert to frequency domain using Short-Time Fourier Transform (STFT)
# n_fft = window size for FFT; hop_length = step between consecutive windows
n_fft = 2048
hop_length = 512
stft_output = librosa.stft(audio_waveform, n_fft=n_fft, hop_length=hop_length)

# Convert complex STFT output to amplitude (magnitude spectrum)
frequency_amplitude_matrix = np.abs(stft_output)

# Get the actual frequency values corresponding to each frequency bin
frequency_bins = librosa.fft_frequencies(sr=sample_rate, n_fft=n_fft)

Quick breakdown:

  • amplitude_array: 1D array where each element is the audio's amplitude at a single time point (length = total audio duration × sample rate)
  • frequency_amplitude_matrix: 2D array (frequency bins × time windows) showing amplitude for each frequency at each time slice
  • frequency_bins: 1D array of the exact frequency values (in Hz) for each row in frequency_amplitude_matrix

Q2: Building a neural network to rate music quality (1-10) from MP3 files

Your core idea is solid, but raw frequency/amplitude arrays are too high-dimensional and noisy for a neural network. Instead, we’ll extract compact, meaningful audio features first, then feed those into a model. Here’s a step-by-step implementation plan:

1. Data Preparation

First, you need a labeled dataset:

  • Collect hundreds/thousands of MP3 files, each tagged with a quality score (1-10). Aim for balanced data (don’t let 90% of files be rated 10!)—use average scores from multiple human reviewers if possible for objectivity.

2. Extract Meaningful Audio Features

Use librosa to extract features that capture musical properties (timbre, pitch, rhythm, etc.). Here’s a function to extract a robust set of features:

import librosa
import numpy as np

def extract_music_features(file_path):
    # Load audio with a fixed sample rate (22050 Hz is standard for librosa)
    y, sr = librosa.load(file_path, sr=22050)
    
    # Mel Spectrogram (captures timbre, aligned with human auditory system)
    mel_spec = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=128)
    mel_spec_db = librosa.power_to_db(mel_spec, ref=np.max)
    
    # MFCCs (Mel-Frequency Cepstral Coefficients—compact timbre representation)
    mfccs = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=13)
    mfccs_delta = librosa.feature.delta(mfccs)  # First-order time difference
    mfccs_delta2 = librosa.feature.delta(mfccs, order=2)  # Second-order difference
    combined_mfccs = np.concatenate([mfccs, mfccs_delta, mfccs_delta2], axis=0)
    
    # Chroma Features (captures pitch class distribution, related to harmony)
    chroma = librosa.feature.chroma_stft(y=y, sr=sr)
    
    # Spectral Contrast (captures difference between spectrum peaks and valleys)
    spectral_contrast = librosa.feature.spectral_contrast(y=y, sr=sr)
    
    # Convert all features to fixed-length vectors (take mean + std for each feature)
    feature_list = []
    for feature in [mel_spec_db, combined_mfccs, chroma, spectral_contrast]:
        feature_list.extend(np.mean(feature, axis=1))
        feature_list.extend(np.std(feature, axis=1))
    
    return np.array(feature_list)

This function outputs a fixed-length vector (regardless of audio duration) that summarizes key musical properties—perfect for a neural network input.

3. Design the Neural Network

Since we’re doing a regression task (predicting a continuous score 1-10), here’s a simple but effective model using Keras/TensorFlow:

from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Dropout
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
import numpy as np

# Assume you have a list of file paths and their corresponding scores
file_paths = ["song1.mp3", "song2.mp3", ...]
scores = [7, 3, ...]

# Extract features for all files
X = np.array([extract_music_features(file) for file in file_paths])
y = np.array(scores)

# Standardize features (critical for neural network performance)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Split into train/validation/test sets
X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, test_size=0.2, random_state=42)
X_train, X_val, y_train, y_val = train_test_split(X_train, y_train, test_size=0.2, random_state=42)

# Build the model
input_shape = (X_scaled.shape[1],)  # Length of your feature vector
model = Sequential([
    Dense(256, activation="relu", input_shape=input_shape),
    Dropout(0.3),  # Prevent overfitting
    Dense(128, activation="relu"),
    Dropout(0.3),
    Dense(64, activation="relu"),
    Dense(1, activation="linear")  # Output continuous score 1-10
])

# Compile the model
model.compile(optimizer="adam", loss="mse", metrics=["mae"])

# Train the model
history = model.fit(
    X_train, y_train,
    validation_data=(X_val, y_val),
    epochs=50,
    batch_size=32
)

4. Evaluate and Deploy

  • After training, test the model on X_test to check prediction accuracy (use MAE—Mean Absolute Error—to measure average prediction error)
  • To rate new MP3s: extract features with extract_music_features(), scale them with the same scaler, then run model.predict(scaled_features) to get the score

Key Tips for Better Performance

  • Use transfer learning: Instead of building from scratch, use pre-trained audio models like YAMNet (trained on millions of audio clips) and fine-tune it on your dataset—this saves time and boosts accuracy
  • Try sequence models: If you want to use time-series data (like the full mel spectrogram), use CNNs (to capture frequency patterns) or LSTMs (to capture temporal patterns)
  • Augment your data: Apply audio augmentations (pitch shifting, time stretching, adding mild noise) to training data to make the model more robust

内容的提问来源于stack exchange,提问作者Dreaded Harvester

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:17:05