用于说话人验证的Siamese网络验证精度无提升问题求助
文本无关说话人验证Siamese网络精度停滞问题
问题背景
开发用于考勤系统的文本无关说话人验证Siamese网络,目标是对比两个音频判断是否为同一说话人,因难以获取每人大量数据采用Siamese架构。流程为加载指定文件夹内的5秒WAV音频片段,构建正负样本对后训练,但模型验证精度始终无提升,仅在0.648-0.729区间波动。
代码文件
process_audio.py(音频加载与预处理)
import os import librosa import numpy as np from tensorflow.image import resize def process_audio(audio_data, sample_rate, target_shape=(128, 128)): length = len(audio_data)/sample_rate max_length = 5*sample_rate if length > 5: audio_data = audio_data[:max_length] mel_spectrogram = librosa.feature.melspectrogram(y=audio_data, sr=sample_rate) mel_spectrogram = resize(np.expand_dims(mel_spectrogram, axis=-1), target_shape) return mel_spectrogram def load_and_preprocess_data(data_dir, classes, target_shape=(128, 128)): data = [] labels = [] for i, class_name in enumerate(classes): class_dir = os.path.join(data_dir, class_name) for filename in os.listdir(class_dir)[:30]: if filename.endswith('.wav'): file_path = os.path.join(class_dir, filename) audio_data, sample_rate = librosa.load(file_path, sr=16000) data.append(process_audio(audio_data, sample_rate, target_shape)) labels.append(i) return np.array(data), np.array(labels) def load_and_process_single(file_path, target_shape=(128, 128)): audio_data, sample_rate = librosa.load(file_path, sr=16000) processed_audio = process_audio(audio_data, sample_rate, target_shape) return np.array([processed_audio])
preprocess_pairs.py(样本对构建)
import numpy as np def load_pairs(data, labels): """ arguments -> data : list of mel spectograms labels : labels from 0 to number of classes returns positive and negative pairs along with their new labels """ groups = [] positive_pairs = [] negative_pairs = [] # split data into groups by label for i in range(len(data)): try: groups[labels[i]].append(data[i]) except: groups.insert(labels[i], []) groups[labels[i]].append(data[i]) # loop through items in the same group and asign them as a positive pair for i in range(len(groups)): for j in range(0, len(groups[i]), 2): pair = (groups[i][j], groups[i][j+1]) positive_pairs.append(pair) # loop through items of different groups and asign negative pairs compared = [] for i in range(0, min(map(len, groups))-1): for j in range(len(groups)): for x in range(len(groups)): if x == j or j in compared: break pair = (groups[j][i], groups[x][i]) negative_pairs.append(pair) compared.append(j) # get rid of the temporary variable del(compared) # comvert to ndarray positive_pairs = np.array(positive_pairs) negative_pairs = np.array(negative_pairs) # concatenate positive and negative pairs into 1 pairs = np.concatenate((positive_pairs, negative_pairs), axis=0) # generate new labels showing difference. 1 means alike, 0 means not alike diffs = np.array([1]*len(positive_pairs) + [0]*len(negative_pairs)) return (pairs, diffs)
train.py(模型训练)
import numpy as np import keras from keras.layers import * from tensorflow.keras.optimizers import * from keras.models import Sequential, Model from sklearn.model_selection import train_test_split from process_audio import * from preprocess_pairs import * audio_data_shape = (128, 128) # List of classes (also folder names) classes = ["Ahmed Hussain", "Alasfoor", "Ali Ayyad", "Ali Jaffar", "Elyas"] data, labels = load_and_preprocess_data("Audio", classes, target_shape=audio_data_shape) # Load data into pairs pairs, diffs = load_pairs(data, labels) # Get train and test splits x_train, x_test, y_train, y_test = train_test_split(pairs, diffs, test_size=0.2, shuffle=True) # Create the cnn model for feature extraction cnn_model = Sequential([ Conv2D(8, (3, 3), activation='relu', input_shape=(128, 128, 1)), BatchNormalization(), Conv2D(16, (3, 3), activation='relu'), BatchNormalization(), GlobalAveragePooling2D(), Dense(64, activation='relu') ]) # Define inputs for compared voices (can be said images as they are melspectograms) input1 = Input(audio_data_shape, name='voice_1') input2 = Input(audio_data_shape, name='voice_2') # convert to feature vector using cnn model feature_vector_1 = cnn_model(input1) feature_vector_2 = cnn_model(input2) # Create the main model with dropout layers to avoid overfitting concat = Concatenate()([feature_vector_1, feature_vector_2]) dropout1 = Dropout(0.5, seed=77)(concat) dense1 = Dense(64, activation='relu')(dropout1) dropout2 = Dropout(0.5, seed=77)(dense1) output_layer2 = Dense(1, activation='sigmoid')(dropout2) model = Model(inputs=[input1, input2], outputs=output_layer2) model.compile(optimizer=Adam(learning_rate=1E-6), loss='binary_crossentropy', metrics=['accuracy']) # Train the model with the extracted data model.fit(x=[x_train[:, 0, :, :], x_train[:, 1, :, :]], y=y_train, epochs=20, batch_size=16, validation_data=[([x_test[:, 0, :, :], x_test[:, 1, :, :]]), y_test])
训练异常表现
训练过程中验证精度始终在0.6-0.7区间波动,无明显提升,部分训练输出如下:
2024-03-25 23:57:04.001297: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: SSE SSE2 SSE3 SSE4.1 SSE4.2 AVX AVX2 AVX_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. C:\Users\NV23142\AppData\Local\Programs\Python\Python311\Lib\site-packages\librosa\core\spectrum.py:257: UserWarning: n_fft=2048 is too large for input signal of length=1920 warnings.warn( C:\Users\NV23142\AppData\Local\Programs\Python\Python311\Lib\site-packages\librosa\core\spectrum.py:257: UserWarning: n_fft=2048 is too large for input signal of length=0 warnings.warn( C:\Users\NV23142\AppData\Local\Programs\Python\Python311\Lib\site-packages\librosa\core\spectrum.py:257: UserWarning: n_fft=2048 is too large for input signal of length=606 warnings.warn( C:\Users\NV23142\AppData\Local\Programs\Python\Python311\Lib\site-packages\librosa\core\spectrum.py:257: UserWarning: n_fft=2048 is too large for input signal of length=286 warnings.warn( Epoch 1/20 13/13 [==============================] - 5s 262ms/step - loss: 0.6883 - accuracy: 0.5743 - val_loss: 0.6769 - val_accuracy: 0.7059 Epoch 2/20 13/13 [==============================] - 3s 239ms/step - loss: 0.6927 - accuracy: 0.5149 - val_loss: 0.6815 - val_accuracy: 0.7059 Epoch 3/20 13/13 [==============================] - 3s 236ms/step - loss: 0.6918 - accuracy: 0.5099 - val_loss: 0.6835 - val_accuracy: 0.7059 Epoch 4/20 13/13 [==============================] - 3s 237ms/step - loss: 0.6930 - accuracy: 0.5198 - val_loss: 0.6845 - val_accuracy: 0.6667 Epoch 5/20 13/13 [==============================] - 3s 237ms/step - loss: 0.6870 - accuracy: 0.5594 - val_loss: 0.6848 - val_accuracy: 0.6471 Epoch 6/20 13/13 [==============================] - 3s 236ms/step - loss: 0.6863 - accuracy: 0.5495 - val_loss: 0.6849 - val_accuracy: 0.6863 Epoch 7/20 13/13 [==============================] - 3s 236ms/step - loss: 0.6902 - accuracy: 0.5446 - val_loss: 0.6849 - val_accuracy: 0.7059 Epoch 8/20 13/13 [==============================] - 3s 245ms/step - loss: 0.6929 - accuracy: 0.5198 - val_loss: 0.6850 - val_accuracy: 0.7255 Epoch 9/20 13/13 [==============================] - 3s 237ms/step - loss: 0.6887 - accuracy: 0.5446 - val_loss: 0.6850 - val_accuracy: 0.7059 Epoch 10/20 13/13 [==============================] - 3s 239ms/step - loss: 0.6909 - accuracy: 0.5495 - val_loss: 0.6851 - val_accuracy: 0.7059 Epoch 11/20 13/13 [==============================] - 3s 239ms/step - loss: 0.6912 - accuracy: 0.5446 - val_loss: 0.6852 - val_accuracy: 0.7059 Epoch 12/20 5/13 [==========>...................] - ETA: 1s - loss: 0.6895 - accuracy: 0.5250Traceback (most recent call last):
已尝试的优化(无改善)
- 降低学习率至1E-6
- 添加Dropout层抑制过拟合
- 降低模型复杂度
解决方案建议
1. 修复音频预处理问题
- 处理短音频:当前代码只截断超过5秒的音频,对不足5秒的音频未做处理,导致部分音频长度过短触发
n_fft警告。修改process_audio函数:
def process_audio(audio_data, sample_rate, target_shape=(128, 128)): max_length = 5 * sample_rate # 补零或截断到固定长度5秒 if len(audio_data) < max_length: audio_data = np.pad(audio_data, (0, max_length - len(audio_data)), mode='constant') else: audio_data = audio_data[:max_length] # 调整n_fft为适合短音频的值,比如512 mel_spectrogram = librosa.feature.melspectrogram(y=audio_data, sr=sample_rate, n_fft=512) mel_spectrogram = resize(np.expand_dims(mel_spectrogram, axis=-1), target_shape) return mel_spectrogram
- 添加音频归一化:对音频数据做标准化(
audio_data = librosa.util.normalize(audio_data)),稳定输入分布。
2. 修正样本对构建逻辑
- 正负样本数量平衡:当前
load_pairs函数的负样本构建逻辑存在问题,可能导致正负样本数量失衡或多样性不足。重新实现配对逻辑:- 正样本:每个说话人的所有音频两两组合(而非步长为2取对)
- 负样本:每个正样本对应等量的不同说话人音频对,确保正负样本数量一致
- 避免重复配对:改用随机采样方式生成负样本,提升多样性。
3. 调整模型与损失函数
- 更换Siamese专用损失:Siamese网络更适合用**对比损失(Contrastive Loss)**而非二元交叉熵,直接优化特征间的距离:
def contrastive_loss(y_true, y_pred): margin = 1 square_pred = K.square(y_pred) margin_square = K.square(K.maximum(margin - y_pred, 0)) return K.mean(y_true * square_pred + (1 - y_true) * margin_square) # 计算特征向量的欧式距离 distance = Lambda(lambda x: K.sqrt(K.sum(K.square(x[0] - x[1]), axis=1, keepdims=True)))([feature_vector_1, feature_vector_2]) model = Model(inputs=[input1, input2], outputs=distance) model.compile(optimizer=Adam(learning_rate=1e-4), loss=contrastive_loss)
- 调整学习率:1E-6的学习率过低,导致模型无法有效更新参数,建议改为1e-4或1e-5,配合
ReduceLROnPlateau实现学习率衰减。 - 优化特征提取网络:当前CNN结构过于简单,可增加卷积层数或改用MaxPooling2D替代GlobalAveragePooling2D,提升特征提取能力。
4. 数据增强
针对音频添加增强操作,提升模型泛化能力:
- 随机添加低强度噪声
- 随机调整语速、音调
- 随机截取5秒长度的音频片段
5. 验证数据集划分
确保训练集和验证集的说话人无重叠,避免数据泄露(按说话人划分数据集,而非随机拆分样本对)。
内容的提问来源于stack exchange,提问作者Mohammed Adel
相关产品推荐
相关产品推荐

