You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决pyannote.audio中的张量尺寸不匹配问题

解决pyannote Speaker Diarization的RuntimeError(张量尺寸不匹配问题)

问题描述

运行pyannote说话人 diarization 代码时触发以下错误:

RuntimeError: Sizes of tensors must match except in dimension 0. Expected size 80000 but got size 79659 for tensor number 11 in the list.

重现代码

from pyannote.audio import Pipeline
from env import hugging_face_token

print('loading pipeline')
pipeline = Pipeline.from_pretrained("pyannote/speaker-diarization",
                                    use_auth_token = hugging_face_token)

print('diarizing')
diarization = pipeline("data/db short intro.ogg")
print('diarizing complete')

运行环境

使用VS Code Dev Container,配置文件如下:

devcontainer.json

// For format details, see https://aka.ms/devcontainer.json. For config options, see the
// README at: https://github.com/devcontainers/templates/tree/main/src/python
{
    "name": "Python 3",
    "image": "mcr.microsoft.com/devcontainers/python:0-3.10",
    "features": {
        "ghcr.io/devcontainers-contrib/features/ffmpeg-apt-get:1": {}
    },
    "postCreateCommand": "pip3 install --user -r requirements.txt"
}

requirements.txt(初始版本)

ipython
torch==1.11.0 
torchvision==0.12.0 
torchaudio==0.11.0 
torchtext==0.12.0
speechbrain==0.5.12
pyannote.audio
git+https://github.com/openai/whisper.git#egg=openai-whisper

已尝试移除依赖版本限制并更新至最新,问题仍存在。


解决方案

1. 强制转换音频格式为模型要求的标准格式

pyannote默认要求输入为16kHz单声道音频,且音频长度需满足模型内部帧处理的对齐要求。使用ffmpeg转换音频:

ffmpeg -i "data/db short intro.ogg" -ar 16000 -ac 1 "data/converted_audio.wav"

替换代码中的音频路径为转换后的文件,重新运行。

2. 手动处理音频张量对齐

如果转换后仍有问题,可手动加载音频并调整张量长度至模型期望的倍数(错误中提到的80000为单批次处理长度):

import torch
import torchaudio
from pyannote.audio import Pipeline
from env import hugging_face_token

# 加载音频
waveform, sample_rate = torchaudio.load("data/db short intro.ogg")

# 转单声道
if waveform.shape[0] > 1:
    waveform = torch.mean(waveform, dim=0, keepdim=True)

# 转16kHz采样率
if sample_rate != 16000:
    resampler = torchaudio.transforms.Resample(sample_rate, 16000)
    waveform = resampler(waveform)

# 补零对齐至80000的倍数
target_length = ((waveform.shape[1] + 79999) // 80000) * 80000
if waveform.shape[1] < target_length:
    waveform = torch.nn.functional.pad(waveform, (0, target_length - waveform.shape[1]))

# 运行pipeline
pipeline = Pipeline.from_pretrained("pyannote/speaker-diarization", use_auth_token=hugging_face_token)
diarization = pipeline({"waveform": waveform, "sample_rate": 16000})

3. 锁定兼容的依赖版本

旧版torch与pyannote存在兼容性问题,更新至稳定兼容版本:
修改requirements.txt为:

ipython
torch>=2.0.0
torchaudio>=2.0.0
speechbrain>=0.5.13
pyannote.audio==2.1.1
git+https://github.com/openai/whisper.git#egg=openai-whisper

重新安装依赖或重建Dev Container。

4. 验证容器内ffmpeg可用性

确保容器内ffmpeg正常工作,执行以下命令检查:

ffmpeg -version

若无法正常输出,需重新安装ffmpeg或检查Dev Container的features配置。


内容的提问来源于stack exchange,提问作者Mr. Enigma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 07:05:20