在torchaudio中如何混合不同长度、采样率的音频张量?
使用torchaudio混合不同采样率的音频完全可行,以下是具体解决方案
问题根源
重采样后长度不一致,本质是原始字节数组的帧数并非严格匹配15秒(比如录音起始/结束的微小误差、字节截取偏差),直接按现有帧数重采样会保留这种差异。但既然原始音频时长都是15秒,我们可以基于这个固定时长对齐重采样后的音频。
方案一:手动对齐后张量相加
先将字节数组转为torch张量,重采样到目标采样率后,通过裁剪/补零对齐到15秒对应的目标长度(16000*15=240000帧,单声道),再完成混合:
import torch import torchaudio import numpy as np # 假设音频为16位PCM格式,将字节数组转为归一化浮点张量 # 若为双声道,需调整reshape逻辑(如.reshape(-1, 2).T) mic_np = np.frombuffer(mic_bytes, dtype=np.int16).astype(np.float32) / 32768.0 speaker_np = np.frombuffer(speaker_bytes, dtype=np.int16).astype(np.float32) / 32768.0 mic_tensor = torch.tensor(mic_np).unsqueeze(0) # 形状: [1, 帧数] speaker_tensor = torch.tensor(speaker_np).unsqueeze(0) # 重采样到16kHz resample_mic = torchaudio.transforms.Resample(44100, 16000) resample_speaker = torchaudio.transforms.Resample(48000, 16000) mic_resampled = resample_mic(mic_tensor) speaker_resampled = resample_speaker(speaker_tensor) # 对齐到15秒的目标帧数 target_frames = 16000 * 15 # 裁剪过长部分,补零填充过短部分 mic_aligned = torch.nn.functional.pad(mic_resampled, (0, max(0, target_frames - mic_resampled.shape[1])))[:, :target_frames] speaker_aligned = torch.nn.functional.pad(speaker_resampled, (0, max(0, target_frames - speaker_resampled.shape[1])))[:, :target_frames] # 混合音频(除以2避免音量过载) mixed_audio = (mic_aligned + speaker_aligned) / 2.0
方案二:用torchaudio调用sox原生混合功能
torchaudio的sox_effects可以直接调用sox的混合能力,无需手动对齐,只需临时保存音频文件作为输入:
import torch import torchaudio import os # 将张量保存为临时wav文件 torchaudio.save("mic_temp.wav", mic_tensor, sample_rate=44100) torchaudio.save("speaker_temp.wav", speaker_tensor, sample_rate=48000) # 构建sox效果链:统一采样率并混合两个音频 effects = [ ["rate", "16000"], # 转16kHz ["mix", "speaker_temp.wav"], # 混合声卡音频 ] # 处理麦克风音频,完成混合 mixed_tensor, mixed_sr = torchaudio.sox_effects.apply_effects_file("mic_temp.wav", effects) # 清理临时文件 os.remove("mic_temp.wav") os.remove("speaker_temp.wav")
注意事项
- 若原始音频是双声道,需调整张量形状(如转为
[2, 帧数]),重采样和对齐逻辑也要对应适配。 - 混合时建议做归一化处理(如除以2),避免音量过载导致的削波失真。
内容的提问来源于stack exchange,提问作者Cheeter_P
相关产品推荐
相关产品推荐

