You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

微调Wav2Vec2时Librosa音频重采样遇参数错误,求解决

解决Librosa Resample参数错误及音频下采样至16kHz方案

错误原因

你遇到的resample() takes 1 positional argument but 3 were given错误,是因为新版Librosa(0.9.0及以上)修改了resample函数的参数结构——不再支持直接传入三个位置参数,必须显式指定orig_sr(原采样率)和target_sr(目标采样率)作为关键字参数。

修复方案(Librosa版本适配)

修改代码中调用librosa.resample的部分,将位置参数改为关键字参数即可解决:

修复后的重采样代码

import librosa
import numpy as np

# List to store new sampling rates
new_sr = []

# Resample each audio signal and store the new sampling rate
for i in range(len(database['psr'])):
    try:
        audio_signal = np.asarray(database['audio'][i])
        original_sr = database['psr'][i]

        if audio_signal.ndim == 1:
            # 显式指定orig_sr和target_sr关键字参数
            resampled_audio = librosa.resample(audio_signal, orig_sr=original_sr, target_sr=16000)
        else:
            resampled_channels = []
            for channel in audio_signal:
                resampled_channel = librosa.resample(channel, orig_sr=original_sr, target_sr=16000)
                resampled_channels.append(resampled_channel)
            resampled_audio = np.array(resampled_channels)

        database['audio'][i] = resampled_audio
        new_sr.append(16000)
    except Exception as e:
        print(f"Error processing audio at index {i}: {e}")

database['newsr'] = new_sr

替代方案:用Torchaudio直接重采样(更适配你的流程)

既然你已经用torchaudio.load加载音频,推荐直接用Torchaudio的resample函数,避免Librosa版本兼容问题,且全程保持张量操作更高效:

import torch
import torchaudio.transforms as T

new_sr = []
resampler = T.Resample(orig_freq=None, new_freq=16000)

for i in range(len(database['psr'])):
    try:
        # 把numpy数组转回张量
        audio_tensor = torch.from_numpy(database['audio'][i]).unsqueeze(0)
        original_sr = database['psr'][i]
        
        # 更新重采样器的原采样率
        resampler.orig_freq = original_sr
        resampled_tensor = resampler(audio_tensor)
        
        # 转回numpy数组存回database
        database['audio'][i] = resampled_tensor.squeeze(0).numpy()
        new_sr.append(16000)
    except Exception as e:
        print(f"Error processing audio at index {i}: {e}")

database['newsr'] = new_sr

额外提示

  • 若必须使用旧版Librosa参数格式,可降级Librosa到0.8.x版本:pip install librosa==0.8.1,但不推荐,新版本修复了更多音频处理问题。
  • Wav2Vec2要求输入为16kHz单声道音频,你的加载代码已经取了speech_array[0],确保了单声道输入,无需额外处理。

内容的提问来源于stack exchange,提问作者SAFIQUL ISLAM UZZAL 203-15-144

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 08:47:38