You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Whisper与pyannote 3.1触发AttributeError: 'list'对象无'get'属性

解决思路

1. 修正说话人分割结果的类型转换

pyannote.audio 3.x版本中,SpeakerSegmentation模型的输出从原来的Annotation对象改为了Segmentation对象,直接调用itertracks会导致返回结构异常。需要先将其转换为Annotation类型:

from pyannote.core import Annotation

# 假设pipeline是初始化好的说话人分割模型
segmentation_result = pipeline(audio_file1)
who_speaks_when1 = segmentation_result.to_annotation()  # 转换为Annotation对象

2. 调整audio.crop的调用逻辑

3.x版本中Audio.crop不再返回(waveform, sample_rate)元组,而是返回一个Audio对象。需要修改代码获取音频数据:

# 替换原有crop调用
cropped_audio = audio.crop(audio_file1, segment)
waveform, sample_rate = cropped_audio.read()  # 从Audio对象读取波形和采样率

3. 确保Whisper输入格式合规

Whisper默认要求输入音频采样率为16kHz,若pyannote返回的音频采样率不匹配,需要转换:

import librosa

# 检查并转换采样率
if sample_rate != 16000:
    waveform = librosa.resample(waveform, orig_sr=sample_rate, target_sr=16000)

# 转写逻辑保持不变
text = model.transcribe(waveform.squeeze())["text"]

完整修正后的代码示例

from pyannote.core import Annotation
import librosa

# 初始化模型(省略原有初始化代码)
# pipeline = SpeakerSegmentation(...)
# model = whisper.load_model(...)

# 处理说话人分割结果
segmentation_result = pipeline(audio_file1)
who_speaks_when1 = segmentation_result.to_annotation()

for segment, _, speaker in who_speaks_when1.itertracks(yield_label=True):
    cropped_audio = audio.crop(audio_file1, segment)
    waveform, sample_rate = cropped_audio.read()
    
    # 统一采样率到16kHz
    if sample_rate != 16000:
        waveform = librosa.resample(waveform, orig_sr=sample_rate, target_sr=16000)
    
    text = model.transcribe(waveform.squeeze())["text"]
    print(f"{segment.start:06.1f}s {segment.end:06.1f}s {speaker}: {text}")

额外排查点

  • 确认pyannote.audio 3.x的依赖版本匹配,避免因依赖冲突导致的异常
  • 打印who_speaks_when1的类型和itertracks返回的元素结构,确认转换后的格式符合预期
  • 检查Whisper版本是否兼容,避免因Whisper自身升级导致的返回格式变化

内容的提问来源于stack exchange,提问作者boredgirl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 03:16:11