使用Torchaudio保存的WAV文件体积为何约为原文件两倍?
问题原因分析
你的猜测完全正确,核心原因是原始WAV文件采用了压缩编码格式,而torchaudio.save()默认以无压缩的32位浮点PCM格式保存音频,导致文件体积翻倍。
具体拆解:
- 原始WAV的本质:WAV是音频容器格式,并非只有无压缩PCM一种编码。你的原文件体积约334KB,对应167226个采样点,计算下来每个采样点约占2字节——这符合16位整数编码(如μ-law、A-law或压缩PCM)的特征,是压缩/低比特深度的表现。
- torchaudio的加载逻辑:
torchaudio.load()会将任何编码的音频统一解码为32位浮点张量(torch.float32),这是它的默认解码规则,加载后的数据已经是无压缩的浮点格式。 - torchaudio的保存逻辑:
torchaudio.save()默认使用无压缩的32位浮点PCM格式保存,每个采样点占用4字节。167226×4≈668904字节,和你保存后的668962字节基本一致(差异来自WAV文件头的少量额外数据)。二次保存时,加载的是无压缩浮点WAV,保存后格式不变,所以体积不再变化。
关键提示与关键词
- WAV容器编码多样性:WAV支持
μ-law、A-law、ADPCM等压缩编码,并非都是无压缩PCM - torchaudio编解码默认行为:
load()默认解码为float32,save()默认保存为32位浮点PCM - 音频比特深度:16位整数(2字节/采样)与32位浮点(4字节/采样)的体积差异
- 原始文件编码验证:可用ffmpeg、Audacity等工具查看原WAV的具体编码格式
你的代码
import os import torchaudio import torch # 原代码漏导入torch,否则torch.equal会报错 ori_audio, ori_sr = torchaudio.load('LJ037-0171.wav') torchaudio.save('LJ037-0171_resave.wav', ori_audio, ori_sr) resave_audio, resave_sr = torchaudio.load('LJ037-0171_resave.wav') print(f'Original sr: {ori_sr}, Resaved sr: {resave_sr}') print(f'Audio tensor equal: {torch.equal(ori_audio, resave_audio)}') print(f'datatype of ori_audio: {ori_audio[0, 1].dtype}') print(f'datatype of resave_audio: {resave_audio[0, 1].dtype}') # 原代码此处变量名笔误 print(f'Shape of ori: {ori_audio.shape}') print(f'Shape of resave: {resave_audio.shape}') # 原代码此处变量名笔误 print(f'File size of original wav: {os.path.getsize("LJ037-0171.wav")}') print(f'File size of resaved wav: {os.path.getsize("LJ037-0171_resave.wav")}')
输出结果
Original sr: 22050, Resaved sr: 22050 Audio tensor equal: True datatype of ori_audio: torch.float32 datatype of resave_audio: torch.float32 Shape of ori: torch.Size([1, 167226]) Shape of resave: torch.Size([1, 167226]) File size of original wav: 334496 File size of resaved wav: 668962
内容的提问来源于stack exchange,提问作者KilinWei
相关产品推荐
相关产品推荐

