You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

加载HDF5数据集到PyTorch DataLoader时遇ValueError报错求助

解决PyTorch DataLoader加载HDF5数据集时的"Invalid dataset identifier"错误

问题情况

我有一个形状为(1000,)的HDF5数据集,单独执行读取代码时能正常输出数据集形状,但将其封装到自定义PyTorch Dataset类并加载到DataLoader时,触发ValueError: Invalid dataset identifier错误,用f.get("default")也会出现相同错误。

单独读取代码

with h5py.File(dataset_path, 'r') as f:
    data = f['default']
    print(data.shape)

输出:

(1000,)

自定义Dataset及DataLoader代码

from torch.utils.data import Dataset, DataLoader
import h5py

# Define the dataset
class MyDataset(Dataset):
    def __init__(self, dataset_path):
        super().__init__()
        with h5py.File(dataset_path, 'r') as f:
            self.data = f['default']

    def __len__(self):
        return len(self.data)

    def __getitem__(self, idx):
        return self.data[idx]

# Load the dataset
dataset_path = 'dataset.hdf5'

train_dataset = MyDataset(dataset_path)
train_loader = DataLoader(train_dataset, shuffle=True)

报错栈

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-80-c3c741b81eff> in <module>
     22 
     23 train_dataset = MyDataset(dataset_path)
---> 24 train_loader = DataLoader(train_dataset, shuffle=True)

6 frames
h5py/_objects.pyx in h5py._objects.with_phil.wrapper()

h5py/_objects.pyx in h5py._objects.with_phil.wrapper()

/usr/local/lib/python3.9/dist-packages/h5py/_hl/dataset.py in shape(self)
    472 
    473         with phil:
---> 474             shape = self.id.shape
    475 
    476         # If the file is read-only, cache the shape to speed-up future uses.

h5py/h5d.pyx in h5py.h5d.DatasetID.shape.__get__()

h5py/h5d.pyx in h5py.h5d.DatasetID.shape.__get__()

h5py/_objects.pyx in h5py._objects.with_phil.wrapper()

h5py/_objects.pyx in h5py._objects.with_phil.wrapper()

h5py/h5d.pyx in h5py.h5d.DatasetID.get_space()

ValueError: Invalid dataset identifier (invalid dataset identifier)

原因分析

问题出在__init__方法中的with语句:h5py的Dataset对象(即f['default'])依赖于打开的文件句柄。当with代码块执行完毕后,文件会被自动关闭,此时self.data就变成了无效的引用,后续调用len(self.data)或self.data[idx]时,就会触发"无效数据集标识符"的错误。

解决方案

有两种常用的修复方式,可根据数据集大小选择:

方法一:加载全部数据到内存(适合小数据集)

在with块内将数据集完整读取为numpy数组,这样数据不再依赖文件句柄,后续操作直接基于内存中的数组进行。

修正后的代码:

from torch.utils.data import Dataset, DataLoader
import h5py

class MyDataset(Dataset):
    def __init__(self, dataset_path):
        super().__init__()
        with h5py.File(dataset_path, 'r') as f:
            # 读取全部数据到内存,转为numpy数组
            self.data = f['default'][()]

    def __len__(self):
        return len(self.data)

    def __getitem__(self, idx):
        return self.data[idx]

# 测试加载
dataset_path = 'dataset.hdf5'
train_dataset = MyDataset(dataset_path)
train_loader = DataLoader(train_dataset, shuffle=True)

for batch in train_loader:
    print(batch.shape)
    break

方法二:保持文件句柄在Dataset生命周期内打开(适合大数据集)

不使用with语句,而是在__init__中直接打开文件并保存句柄,在Dataset对象被销毁时关闭文件,确保Dataset的整个生命周期内文件都是打开状态。

修正后的代码:

from torch.utils.data import Dataset, DataLoader
import h5py

class MyDataset(Dataset):
    def __init__(self, dataset_path):
        super().__init__()
        # 直接打开文件并保存句柄
        self.f = h5py.File(dataset_path, 'r')
        self.data = self.f['default']

    def __len__(self):
        return len(self.data)

    def __getitem__(self, idx):
        return self.data[idx]

    def __del__(self):
        # 对象销毁时关闭文件
        self.f.close()

# 测试加载
dataset_path = 'dataset.hdf5'
train_dataset = MyDataset(dataset_path)
train_loader = DataLoader(train_dataset, shuffle=True)

for batch in train_loader:
    print(batch.shape)
    break

内容的提问来源于stack exchange,提问作者daylightisminetocommand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 22:57:56