使用Whisper与pyannote 3.1触发AttributeError: 'list'对象无'get'属性
解决思路
1. 修正说话人分割结果的类型转换
pyannote.audio 3.x版本中,SpeakerSegmentation模型的输出从原来的Annotation对象改为了Segmentation对象,直接调用itertracks会导致返回结构异常。需要先将其转换为Annotation类型:
from pyannote.core import Annotation # 假设pipeline是初始化好的说话人分割模型 segmentation_result = pipeline(audio_file1) who_speaks_when1 = segmentation_result.to_annotation() # 转换为Annotation对象
2. 调整audio.crop的调用逻辑
3.x版本中Audio.crop不再返回(waveform, sample_rate)元组,而是返回一个Audio对象。需要修改代码获取音频数据:
# 替换原有crop调用 cropped_audio = audio.crop(audio_file1, segment) waveform, sample_rate = cropped_audio.read() # 从Audio对象读取波形和采样率
3. 确保Whisper输入格式合规
Whisper默认要求输入音频采样率为16kHz,若pyannote返回的音频采样率不匹配,需要转换:
import librosa # 检查并转换采样率 if sample_rate != 16000: waveform = librosa.resample(waveform, orig_sr=sample_rate, target_sr=16000) # 转写逻辑保持不变 text = model.transcribe(waveform.squeeze())["text"]
完整修正后的代码示例
from pyannote.core import Annotation import librosa # 初始化模型(省略原有初始化代码) # pipeline = SpeakerSegmentation(...) # model = whisper.load_model(...) # 处理说话人分割结果 segmentation_result = pipeline(audio_file1) who_speaks_when1 = segmentation_result.to_annotation() for segment, _, speaker in who_speaks_when1.itertracks(yield_label=True): cropped_audio = audio.crop(audio_file1, segment) waveform, sample_rate = cropped_audio.read() # 统一采样率到16kHz if sample_rate != 16000: waveform = librosa.resample(waveform, orig_sr=sample_rate, target_sr=16000) text = model.transcribe(waveform.squeeze())["text"] print(f"{segment.start:06.1f}s {segment.end:06.1f}s {speaker}: {text}")
额外排查点
- 确认pyannote.audio 3.x的依赖版本匹配,避免因依赖冲突导致的异常
- 打印
who_speaks_when1的类型和itertracks返回的元素结构,确认转换后的格式符合预期 - 检查Whisper版本是否兼容,避免因Whisper自身升级导致的返回格式变化
内容的提问来源于stack exchange,提问作者boredgirl
相关产品推荐
相关产品推荐

