You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用np.array()转换频谱图数组出现广播错误的解决咨询

问题解决方案

问题根源

你的代码错误地对频谱图的mel特征维度(第一个维度,64)做了填充/截断,而非统一时间维度(第二个维度,如400、469),导致每个样本的时间维度长度不一致,无法合并为统一的numpy数组,触发广播错误。


具体解决方案

1. 修正load_spectrogram函数

重新调整逻辑,针对时间维度做统一的截断/填充,而非mel维度:

def load_spectrogram(file_path, target_time_steps=500):
    # 直接指定采样率为16000,避免重复配置
    audio, sr = librosa.load(file_path, sr=16000)
    window_size_sec = 0.025
    hop_size_sec = 0.01

    window_size = int(16000 * window_size_sec)
    hop_size = int(16000 * hop_size_sec)

    n_fft = 2 ** int(np.ceil(np.log2(window_size)))  

    # 生成mel频谱图
    spectrogram = librosa.feature.melspectrogram(
        y=audio, sr=16000, n_fft=n_fft, hop_length=hop_size, n_mels=64
    )
    spectrogram = librosa.power_to_db(spectrogram, ref=np.max)
    spectrogram = spectrogram.astype(np.float32)
    spectrogram = np.expand_dims(spectrogram, axis=-1)  # 形状:(64, T, 1)
    
    # 统一时间维度:截断过长样本,填充过短样本
    current_time_steps = spectrogram.shape[1]
    if current_time_steps > target_time_steps:
        spectrogram_processed = spectrogram[:, :target_time_steps, :]
    else:
        # 在时间维度后补0,保持mel维度和通道维度不变
        pad_width = ((0, 0), (0, target_time_steps - current_time_steps), (0, 0))
        spectrogram_processed = np.pad(spectrogram, pad_width, mode='constant', constant_values=0)
    
    print(f'Processed spectrogram shape: {spectrogram_processed.shape}')
    return spectrogram_processed

2. 调整数据集加载流程

如果不想固定时间步长,可先遍历所有音频获取全局最大时间步长,再统一处理:

# 加载音频路径与标签
Audios = []
tkinter.Tk().withdraw()
folder_path = filedialog.askdirectory()

for x in tqdm(os.listdir(folder_path)):
    sub_folder = os.path.join(folder_path, x)
    if not os.path.isdir(sub_folder):
        continue  # 跳过非文件夹文件
    for y in os.listdir(sub_folder):
        file_path = os.path.join(sub_folder, y)
        Audios.append(Audio(file_path, label=x))  # 直接用子文件夹名作为标签

# 第一步:计算所有音频的最大时间步长
max_time_steps = 0
for audio in Audios:
    audio_data, sr = librosa.load(audio.file_path, sr=16000)
    window_size = int(16000 * 0.025)
    hop_size = int(16000 * 0.01)
    n_fft = 2 ** int(np.ceil(np.log2(window_size)))
    time_steps = (audio_data.shape[0] - n_fft) // hop_size + 1
    max_time_steps = max(max_time_steps, time_steps)

print(f'全局最大时间步长: {max_time_steps}')

# 第二步:加载并统一所有频谱图
X = []
y = []
for audio in Audios:
    spectrogram = load_spectrogram(audio.file_path, target_time_steps=max_time_steps)
    normalized_spectrogram = normalize_spectrogram(spectrogram)
    X.append(normalized_spectrogram)
    y.append(audio.label)

# 现在可顺利转为numpy数组
X = np.array(X)
y = np.array(y)

# 拆分数据集
X_train, y_train, X_test, y_test = split_data(X, y)

关键改动说明

  • 移除了错误的pad_sequences误用:原代码对mel特征维度(第一个维度)填充,改为针对时间维度(第二个维度)用np.pad手动处理。
  • 删除了逻辑矛盾的spectrogram_padded[:time_steps]截断操作,避免破坏维度统一。
  • 简化标签获取逻辑,直接用子文件夹名作为标签,减少路径拆分错误。
  • 提前统一采样率配置,避免重复定义变量。

额外建议

  • 若模型支持可变长度输入,可使用tf.data.Dataset加载数据,无需提前统一时间维度,适合长音频场景。
  • 固定target_time_steps时,建议根据音频时长分布选择合理值(如覆盖95%以上样本的时间步长),避免过度截断或填充。

内容的提问来源于stack exchange,提问作者CitrusBoy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 10:14:56