加载HDF5数据集到PyTorch DataLoader时遇ValueError报错求助
解决PyTorch DataLoader加载HDF5数据集时的"Invalid dataset identifier"错误
问题情况
我有一个形状为(1000,)的HDF5数据集,单独执行读取代码时能正常输出数据集形状,但将其封装到自定义PyTorch Dataset类并加载到DataLoader时,触发ValueError: Invalid dataset identifier错误,用f.get("default")也会出现相同错误。
单独读取代码
with h5py.File(dataset_path, 'r') as f: data = f['default'] print(data.shape)
输出:
(1000,)
自定义Dataset及DataLoader代码
from torch.utils.data import Dataset, DataLoader import h5py # Define the dataset class MyDataset(Dataset): def __init__(self, dataset_path): super().__init__() with h5py.File(dataset_path, 'r') as f: self.data = f['default'] def __len__(self): return len(self.data) def __getitem__(self, idx): return self.data[idx] # Load the dataset dataset_path = 'dataset.hdf5' train_dataset = MyDataset(dataset_path) train_loader = DataLoader(train_dataset, shuffle=True)
报错栈
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) <ipython-input-80-c3c741b81eff> in <module> 22 23 train_dataset = MyDataset(dataset_path) ---> 24 train_loader = DataLoader(train_dataset, shuffle=True) 6 frames h5py/_objects.pyx in h5py._objects.with_phil.wrapper() h5py/_objects.pyx in h5py._objects.with_phil.wrapper() /usr/local/lib/python3.9/dist-packages/h5py/_hl/dataset.py in shape(self) 472 473 with phil: ---> 474 shape = self.id.shape 475 476 # If the file is read-only, cache the shape to speed-up future uses. h5py/h5d.pyx in h5py.h5d.DatasetID.shape.__get__() h5py/h5d.pyx in h5py.h5d.DatasetID.shape.__get__() h5py/_objects.pyx in h5py._objects.with_phil.wrapper() h5py/_objects.pyx in h5py._objects.with_phil.wrapper() h5py/h5d.pyx in h5py.h5d.DatasetID.get_space() ValueError: Invalid dataset identifier (invalid dataset identifier)
原因分析
问题出在__init__方法中的with语句:h5py的Dataset对象(即f['default'])依赖于打开的文件句柄。当with代码块执行完毕后,文件会被自动关闭,此时self.data就变成了无效的引用,后续调用len(self.data)或self.data[idx]时,就会触发"无效数据集标识符"的错误。
解决方案
有两种常用的修复方式,可根据数据集大小选择:
方法一:加载全部数据到内存(适合小数据集)
在with块内将数据集完整读取为numpy数组,这样数据不再依赖文件句柄,后续操作直接基于内存中的数组进行。
修正后的代码:
from torch.utils.data import Dataset, DataLoader import h5py class MyDataset(Dataset): def __init__(self, dataset_path): super().__init__() with h5py.File(dataset_path, 'r') as f: # 读取全部数据到内存,转为numpy数组 self.data = f['default'][()] def __len__(self): return len(self.data) def __getitem__(self, idx): return self.data[idx] # 测试加载 dataset_path = 'dataset.hdf5' train_dataset = MyDataset(dataset_path) train_loader = DataLoader(train_dataset, shuffle=True) for batch in train_loader: print(batch.shape) break
方法二:保持文件句柄在Dataset生命周期内打开(适合大数据集)
不使用with语句,而是在__init__中直接打开文件并保存句柄,在Dataset对象被销毁时关闭文件,确保Dataset的整个生命周期内文件都是打开状态。
修正后的代码:
from torch.utils.data import Dataset, DataLoader import h5py class MyDataset(Dataset): def __init__(self, dataset_path): super().__init__() # 直接打开文件并保存句柄 self.f = h5py.File(dataset_path, 'r') self.data = self.f['default'] def __len__(self): return len(self.data) def __getitem__(self, idx): return self.data[idx] def __del__(self): # 对象销毁时关闭文件 self.f.close() # 测试加载 dataset_path = 'dataset.hdf5' train_dataset = MyDataset(dataset_path) train_loader = DataLoader(train_dataset, shuffle=True) for batch in train_loader: print(batch.shape) break
内容的提问来源于stack exchange,提问作者daylightisminetocommand
相关产品推荐
相关产品推荐

