大HDF5文件中访问数据集shape速度骤降45倍的问题排查与优化方案求助
大HDF5文件中访问数据集shape速度骤降45倍的问题排查与优化方案求助
我在递归访问一个包含大量数据集的超大HDF5文件时,遇到了读取速度急剧下降的问题,想请大家帮忙分析原因并给出优化建议!
问题背景
我有两个HDF5文件:
small.hdf5:大小122GB,包含119,189个数据集,遍历速度约9000it/slarge.hdf5:大小1.5TB,包含1,000,416个数据集,遍历速度仅约200it/s,速度直接下降了45倍
用cProfile性能分析后发现,__getitem__方法的耗时随文件中组数量增加暴增,尤其是执行self.__file[subject]['eeg'].shape[1]这行代码的时候。
测试代码与性能分析
核心测试代码
self.__file = h5py.File(str(self.__file_path), 'r',) self.__subjects = [i for i in self.__file] import cProfile import pstats import io profiler = cProfile.Profile() profiler.enable() ssum = 0 for subject in tqdm(self.__subjects, desc="Processing subjects (SingleShockDataset)"): subject_len = self.__file[subject]['eeg'].shape[1] # 为了提前终止,避免耗时太长 ssum += 1 if ssum > 10000: break profiler.disable() s = io.StringIO() sortby = 'cumulative' ps = pstats.Stats(profiler, stream=s).sort_stats(sortby) ps.print_stats(10) print(s.getvalue())
针对small.hdf5的性能分析结果
ncalls tottime percall cumtime percall filename:lineno(function) 2/1 0.000 0.000 1.292 1.292 /global/common/software/m4244/DIVER/lib/python3.12/threading.py:637(wait) 2/1 0.000 0.000 1.292 1.292 /global/common/software/m4244/DIVER/lib/python3.12/threading.py:323(wait) 9/3 0.168 0.019 1.292 0.431 {method 'acquire' of '_thread.lock' objects} 20002 0.598 0.000 1.047 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/group.py:348(__getitem__) 10001 0.185 0.000 0.208 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/dataset.py:659(__init__) 10001 0.024 0.000 0.129 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/base.py:278(file) 10001 0.052 0.000 0.091 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/files.py:376(__init__) 10001 0.066 0.000 0.067 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/dataset.py:485(shape) 60007 0.029 0.000 0.044 0.000 <frozen importlib._bootstrap>:1390(_handle_fromlist) 60007 0.021 0.000 0.034 0.000 <frozen importlib._bootstrap>:645(parent)
针对large.hdf5的性能分析结果
ncalls tottime percall cumtime percall filename:lineno(function) 7 0.101 0.014 82.463 11.780 /global/common/software/m4244/DIVER/lib/python3.12/threading.py:637(wait) 20002 65.130 0.003 66.293 0.003 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/group.py:348(__getitem__) 7 0.013 0.002 62.463 8.923 /global/common/software/m4244/DIVER/lib/python3.12/threading.py:323(wait) 28 0.214 0.008 42.437 1.516 {method 'acquire' of '_thread.lock' objects} 10001 0.596 0.000 0.657 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/dataset.py:659(__init__) 10001 0.059 0.000 0.238 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/base.py:278(file) 10002 0.028 0.000 0.198 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/tqdm/std.py:1160(__iter__) 659 0.005 0.000 0.168 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/tqdm/std.py:1198(update) 10001 0.164 0.000 0.165 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/dataset.py:485(shape) 660 0.003 0.000 0.159 0.000 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/tqdm/std.py:1325(refresh)
已尝试的优化方法(均未解决问题)
方法1:直接遍历组的items()
我以为提前拿到组对象可以避免重复搜索,但速度依然很慢:
for group_name, group in tqdm(self.__file.items(), desc="Processing groups (SingleShockDataset)"): subject_len = group['eeg'].shape[1]
对应的性能分析核心结果:
ncalls tottime percall cumtime percall filename:lineno(function) 20002 62.178 0.003 63.334 0.003 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/group.py:348(__getitem__)
方法2:分批处理数据集
尝试批量加载数据集,速度几乎没有变化:
def _collect_eeg_datasets(self): eeg_datasets = [] def visitor(name, obj): if isinstance(obj, h5py.Dataset) and name.endswith('eeg'): eeg_datasets.append(obj) self.__file.visititems(visitor) return eeg_datasets # 主逻辑中的分批处理代码 dataset_names = list(self.__file.keys()) num_datasets = len(dataset_names) batch_size = 50 for i in tqdm(range(0, num_datasets, batch_size), desc="Processing EEG datasets in Batches"): batch_names = dataset_names[i:i + batch_size] # 仅访问数据集对象,不加载数据到内存 eeg_datasets = [self.__file[name]['eeg'] for name in batch_names] for eeg_dataset in eeg_datasets: subject_len = eeg_dataset.shape[1] ssum += 1 if ssum > 10000: break
对应的性能分析核心结果:
ncalls tottime percall cumtime percall filename:lineno(function) 20100 40.024 0.002 41.035 0.002 /global/common/software/m4244/DIVER/lib/python3.12/site-packages/h5py/_hl/group.py:348(__getitem__)
我的疑问
我对编程了解不多,想请教大家:
- 我的评估是否正确?是不是HDF5文件中组的数量太多,导致h5py每次查询都要花费大量时间搜索?
- 有没有可行的优化方案?比如直接遍历数据集而不是按顺序查询?
- 还有没有其他可能的原因或解决思路?
提前感谢大家的建议和评论!
备注:内容来源于stack exchange,提问作者Danny Han
相关产品推荐
相关产品推荐

