You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch中用无梯度损失训练音频增强Transformer模型的可行性

音频质量提升模型训练梯度问题

我正在开发一款音频质量提升模型,采用DAC构建latent space,流程为:含噪音频→编码→latent space A→待训练Transformer→latent space B→解码→降噪音频。

我希望基于原始音频训练Transformer:训练时解码生成音频,计算原始音频与解码后音频的L1损失。但解码过程无法保留梯度(内存占用过高,且解码器无需训练)。我疑惑PyTorch为何需要解码器梯度来更新Transformer权重,按链式法则,梯度应该在Transformer输出处终止。

当前报错信息

loss.backward()
File "/home/jourdelune/Bureau/dev/WaveAI/AudioEnhancer/venv/lib/python3.10/site-packages/torch/_tensor.py", line 492, in backward
    torch.autograd.backward(
File "/home/jourdelune/Bureau/dev/WaveAI/AudioEnhancer/venv/lib/python3.10/site-packages/torch/autograd/__init__.py", line 251, in backward
    Variable._execution_engine.run_backward(  # Calls into the C++ engine to run the backward pass
RuntimeError: element 0 of tensors does not require grad and does not have a grad_fn

期望的训练代码

y_hat = model(x)  # prediction of the model
y_hat = rearrange(y_hat, "b (t c) d -> b c d t", c=c, d=d)  # transform the data to have the latent space dim
loss = 0
for i in range(y_hat.shape[0]):  # for each value in the batch, DAC doesn't support decoding of batches
    with torch.no_grad():  # disable gradient because too much RAM
        z_q, _, _, _ = dataset.autoencoder.quantizer(y_hat[i].float(), None)
        decoded = dataset.autoencoder.decode(z_q)  # RAW audio
    out = l1_loss(decoded, base_waveform[i][:, :, : decoded.shape[-1]])  # compute the distance between the base waveform (target) and the decoded audio (pred)
    loss += out
loss.backward()  # backpropagation

尝试过的写法

loss = torch.tensor(loss, requires_grad=True)
loss.backward()

这种训练方式是否可行?

内容的提问来源于stack exchange,提问作者Jourdelune

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 09:50:04