深度学习入门求助:音频分类模型预测输入形状不匹配问题
model.predict() Hey there! That shape mismatch error is super common when you're starting out with Keras—let's break down what's going on and fix it quickly.
The Root Problem
Your model was trained on input data with shape (number_of_samples, 40) (each sample is a 40-dimensional MFCC vector). When you call model.predict(), Keras expects inputs to still follow this batch-first format: the first dimension is the number of samples you're predicting at once (the batch size), and the second is the feature count (40).
In your test code, mfccs ends up as a 1-dimensional array with shape (40,). When you pass this directly to predict(), Keras misinterprets it as a batch of 40 samples, each with 1 feature—hence the error saying it expected (None, 40) but got (40, 1).
The Fix: Add a Batch Dimension
You just need to convert your single sample into a "batch of 1" by adding an extra dimension. There are two easy ways to do this:
Option 1: Use np.expand_dims()
This is explicit and clear:
file_name = ".../UrbanSoundClassifier/test/Test/5.wav" test_X, sample_rate = librosa.load(file_name,res_type='kaiser_fast') mfccs = np.mean(librosa.feature.mfcc(y=test_X, sr=sample_rate, n_mfcc=40).T,axis=0) # Add a batch dimension at index 0 test_X = np.expand_dims(mfccs, axis=0) # Now shape is (1, 40) print(model.predict(test_X))
Option 2: Reshape the array
You can also use reshape() to force the correct shape:
test_X = mfccs.reshape(1, -1) # The -1 tells numpy to infer the remaining dimension (40 here)
Bonus: Get the Predicted Class
If you want to see which class the model predicts instead of just the raw probability array, use np.argmax():
predictions = model.predict(test_X) predicted_class_index = np.argmax(predictions, axis=1)[0] print(f"Predicted class index: {predicted_class_index}")
For Multiple Samples
If you ever want to predict on multiple audio files at once, just stack all their MFCC vectors into a 2D array with shape (number_of_test_samples, 40)—then you can pass it directly to model.predict() without modifying each one individually.
内容的提问来源于stack exchange,提问作者user3754646

