Keras训练音频分类CNN报错:无法将NumPy数组转为Tensor(不支持list类型)
问题根源
- 核心原因1:生成的2.5秒MFCC片段形状不统一。你没有提前统一所有音频的采样率,且
librosa.time_to_frames计算的帧索引受采样率、hop_length参数影响,不同音频的2.5秒片段对应的MFCC时间维度长度可能存在±1的偏差,所有MFCC特征的shape不是完全一致的(n_mfcc, 固定帧数),转numpy数组时会生成嵌套的object类型数组,无法转换为Tensor。 - 核心原因2:加载数据集时直接把嵌套列表转成
dtype=object的数组,TensorFlow不支持该类型数组作为模型输入。 - 额外错误:构建模型的输入形状定义错误,
input_shape = (X_train.shape[0], X_train.shape[1], X_train.shape[2])把样本数也放入了输入维度,CNN的输入形状不需要带样本数维度。
解决步骤
1. 特征提取阶段统一MFCC长度
首先提前把所有音频统一重采样到固定采样率(比如16000Hz),不要使用原音频的采样率,再提前计算2.5秒对应的固定帧数:
fixed_sample_rate = 16000 fixed_frames = librosa.time_to_frames(2.5, sr=fixed_sample_rate, hop_length=hop_size)
切分出2.5秒的MFCC块后,强制对齐到固定长度:
# 对切出的mfcc_split_in_blocks做对齐,超长截断,不足补0 if mfcc_split_in_blocks.shape[1] > fixed_frames: mfcc_split_in_blocks = mfcc_split_in_blocks[:, :fixed_frames] elif mfcc_split_in_blocks.shape[1] < fixed_frames: pad_len = fixed_frames - mfcc_split_in_blocks.shape[1] mfcc_split_in_blocks = np.pad(mfcc_split_in_blocks, ((0,0), (0,pad_len)), mode='constant')
2. 修正数据集加载代码
现在所有MFCC特征形状完全统一,加载时不需要使用dtype=object,直接转成float32类型的numpy数组即可:
def load_dataset(data_path): list_data_X = [] list_data_y = [] files = [f for f in os.listdir(data_path) if os.path.isfile(os.path.join(data_path, f))] for f in files: path_to_json = os.path.join(data_path, f) with open(path_to_json, "r") as fp: data = json.load(fp) X = np.array(data["mfcc"], dtype=np.float32) y = int(data["label"]) list_data_X.append(X) list_data_y.append(y) X_arr = np.array(list_data_X, dtype = np.float32) y_arr = np.array(list_data_y, dtype = np.int32) return X_arr, y_arr
3. 修正模型输入形状
X_train加通道维度后的形状为(样本数, n_mfcc, 固定帧数, 1),输入形状不需要包含样本维度,修正为:
input_shape = (X_train.shape[1], X_train.shape[2], 1)
4. 可选优化
不要将每个样本单独存为json文件,IO效率低且容易出错,可以直接将所有特征和标签存为npz格式文件,读写速度更快、占用空间更小:
# 保存特征 np.savez('mfcc_features.npz', X=X_arr, y=y_arr) # 加载特征 data = np.load('mfcc_features.npz') X = data['X'] y = data['y']
内容的提问来源于stack exchange,提问作者user16345050
相关产品推荐
相关产品推荐

