基于文本的AI音乐生成模型训练数据基数歧义及文本全为1问题
文本驱动AI音乐生成模型问题排查与解决方案
问题描述
开发基于文本的AI音乐生成模型时,出现两个核心问题:
- 训练时输入的文本样本(x)全部为1
- 触发数据维度不匹配错误:
ValueError: Data cardinality is ambiguous. Make sure all arrays contain the same number of samples. 'x' sizes: 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1 'y' sizes: 3982, 9838, 7699, 14171, 8417, 6547, 9082, 6763, 10255, 11647, 6754, 9491, 17134, 8414, 10078, 8303, 9107, 14320, 10145, 7339, 10338, 6933, 9302, 11573, 11859, 8915, 8949, 9709, 10538, 10924, 9507, 9987, 9445, 11571, 9563, 9705, 10081, 11245, 10239, 11909, 9678, 8788, 10041
核心问题分析
- 文本输入全为1:
- Tokenizer未配置未登录词(OOV)处理,若prompt中词汇未被正确收录,会生成空序列,填充后可能出现异常默认值;额外的维度扩展操作导致文本样本维度混乱,模型误判输入数据。
- 数据维度不匹配:
- 文本样本维度错误,每个样本被处理为
(1,1,43)的数组,Keras误将每个数组视为独立样本,导致样本数与音频数据不匹配。 - 音频MFCC数据长度不一致(不同音频的时间步不同),且模型结构错误:当前模型输出为固定维度向量,无法对应可变长度的MFCC序列。
- 文本样本维度错误,每个样本被处理为
具体修复步骤
一、修复文本输入问题
配置Tokenizer的未登录词处理
修改Tokenizer初始化代码,添加oov_token参数,避免未收录词汇被忽略:tokenizer = Tokenizer(num_words=5000, oov_token="<OOV>")修正文本序列维度
移除不必要的维度扩展,确保每个文本样本为(max_text_len,)的一维数组:def preprocess_data(row, tokenizer, music_dir): song_path = os.path.join(music_dir, row['song']) print(f'Retrieved song path: {song_path}') audio, sr = load_audio(song_path) mfccs = extract_features(audio, sr) text = row['prompt'] sequence = encode_text(text, tokenizer) # 取pad_sequences结果的第一个元素,得到(43,)的数组 padded_sequence = pad_sequences([sequence], maxlen=43, padding='post')[0] return mfccs, padded_sequence转换文本数据为统一格式
将预处理后的文本列表转换为numpy数组,确保维度为(样本数, max_text_len):preprocessed_text = np.array(preprocessed_text)
二、修复音频数据与模型结构
统一音频MFCC时间步长度
对所有MFCC数据进行填充,确保长度一致:# 计算所有MFCC的最大时间步长度 max_mfcc_timesteps = max([mfcc.shape[0] for mfcc in preprocessed_audio]) # 对MFCC进行填充,统一维度 preprocessed_audio = pad_sequences( preprocessed_audio, maxlen=max_mfcc_timesteps, padding='post', dtype='float32' )重构Seq2Seq模型适配音乐生成
修改模型结构,使其输出与MFCC序列匹配的时间分布式结果:max_text_len = 43 mfcc_features = example_mfccs.shape[1] # 获取MFCC的特征维度 num_lstm_units = 128 text_input = Input(shape=(max_text_len,), name='text_input') # 文本编码器:将文本序列转换为上下文向量 text_embedding = Embedding(input_dim=tokenizer.num_words, output_dim=64)(text_input) text_encoder = LSTM(num_lstm_units, return_sequences=False)(text_embedding) # 解码器:将上下文向量重复为MFCC时间步长度,生成序列输出 decoder_input = tf.keras.layers.RepeatVector(max_mfcc_timesteps)(text_encoder) decoder = LSTM(num_lstm_units, return_sequences=True)(decoder_input) # 时间分布式层:输出每个时间步的MFCC特征 decoder_output = TimeDistributed(Dense(mfcc_features))(decoder) model_output = Activation('sigmoid', name='model_output')(decoder_output) model = Model(inputs=text_input, outputs=model_output) model.compile(loss='mse', optimizer='adam')
三、额外排查步骤
- 检查Tokenizer的词汇映射是否正确:
print(tokenizer.word_index) - 打印单个prompt的编码结果,验证是否符合预期:
test_prompt = all_prompts[0] print(f"Prompt: {test_prompt}") print(f"Encoded sequence: {encode_text(test_prompt, tokenizer)}")
内容的提问来源于stack exchange,提问作者progremer
相关产品推荐
相关产品推荐

