DataLoader迭代报错:TypeError与KeyError问题排查求助
问题分析
- TypeError 根源:
__getitem__返回的filename是Path对象,PyTorch默认的default_collate函数无法处理该类型(仅支持张量、numpy数组、数字、字典、列表),因此抛出类型错误。 - KeyError: 0 根源:自定义
collate方法返回的是元组(images, slide_ids, filenames),但遍历代码却用字典索引data['image']取值;同时fastai的错误处理逻辑在捕获TypeError后,错误地将元组当作字典去索引键0,最终引发KeyError。 - 隐藏问题:虽然
DataLoader指定了collate_fn,但报错显示default_collate仍被调用,推测是使用了fastai的DataLoader而非PyTorch原生实现,导致自定义collate逻辑未生效。
解决方案
方案1:统一返回字典格式(适配原遍历代码)
修改collate方法返回字典,同时提前将Path转为字符串避免类型错误:
# 修改ApplicationDataset的__getitem__方法 def __getitem__(self, idx): image = read_image(str(self.tile_filenames[idx])) return { 'image': image, 'slide_id': self.slide_ids[idx], 'filename': str(self.tile_filenames[idx]), # 直接转为字符串 } # 修改collate方法返回字典 @staticmethod def collate(batch): images = torch.stack([batch_item['image'] for batch_item in batch], dim=0) slide_ids = [batch_item['slide_id'] for batch_item in batch] # 字符串类型不能转为tensor,保留列表 filenames = [batch_item['filename'] for batch_item in batch] return { 'image': images, 'slide_id': slide_ids, 'filename': filenames }
原遍历代码无需修改,可正常通过键取值。
方案2:调整遍历代码适配元组返回格式
若要保持collate返回元组,需修改遍历代码并修复slide_ids的类型错误:
# 修改collate方法 @staticmethod def collate(batch): images = torch.stack([batch_item['image'] for batch_item in batch], dim=0) slide_ids = [batch_item['slide_id'] for batch_item in batch] # 字符串列表,避免转tensor报错 filenames = [str(batch_item['filename']) for batch_item in batch] # Path转字符串 return images, slide_ids, filenames # 修改遍历代码 for data in dataloader: image, slide_ids, filenames = data # 直接解包元组 # predict
额外建议
若使用fastai的DataLoader,建议切换为PyTorch原生的torch.utils.data.DataLoader,确保自定义collate_fn被正确执行;若必须使用fastai加载器,需按照其文档规范配置collate逻辑。
内容的提问来源于stack exchange,提问作者DanielBell99
相关产品推荐
相关产品推荐

