PyTorch中用无梯度损失训练音频增强Transformer模型的可行性
音频质量提升模型训练梯度问题
我正在开发一款音频质量提升模型,采用DAC构建latent space,流程为:含噪音频→编码→latent space A→待训练Transformer→latent space B→解码→降噪音频。
我希望基于原始音频训练Transformer:训练时解码生成音频,计算原始音频与解码后音频的L1损失。但解码过程无法保留梯度(内存占用过高,且解码器无需训练)。我疑惑PyTorch为何需要解码器梯度来更新Transformer权重,按链式法则,梯度应该在Transformer输出处终止。
当前报错信息
loss.backward() File "/home/jourdelune/Bureau/dev/WaveAI/AudioEnhancer/venv/lib/python3.10/site-packages/torch/_tensor.py", line 492, in backward torch.autograd.backward( File "/home/jourdelune/Bureau/dev/WaveAI/AudioEnhancer/venv/lib/python3.10/site-packages/torch/autograd/__init__.py", line 251, in backward Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass RuntimeError: element 0 of tensors does not require grad and does not have a grad_fn
期望的训练代码
y_hat = model(x) # prediction of the model y_hat = rearrange(y_hat, "b (t c) d -> b c d t", c=c, d=d) # transform the data to have the latent space dim loss = 0 for i in range(y_hat.shape[0]): # for each value in the batch, DAC doesn't support decoding of batches with torch.no_grad(): # disable gradient because too much RAM z_q, _, _, _ = dataset.autoencoder.quantizer(y_hat[i].float(), None) decoded = dataset.autoencoder.decode(z_q) # RAW audio out = l1_loss(decoded, base_waveform[i][:, :, : decoded.shape[-1]]) # compute the distance between the base waveform (target) and the decoded audio (pred) loss += out loss.backward() # backpropagation
尝试过的写法
loss = torch.tensor(loss, requires_grad=True) loss.backward()
这种训练方式是否可行?
内容的提问来源于stack exchange,提问作者Jourdelune
相关产品推荐
相关产品推荐

