Fortran中无需大内存数组向HDF5文件追加写入数据的实现方法
Fortran HDF5逐次追加写入数据集实现方法
核心原理
要实现类似ASCII文件逐次写入、无需全量内存数组的效果,需要使用HDF5的可扩展数据集功能:
- 创建数据集时将维度的最大值设为
H5S_UNLIMITED_F,允许后续动态修改数据集尺寸 - 每次写入数据前先扩展数据集大小,匹配本次要写入的内容长度
- 通过 hyperslab 选中数据集内新增的空白区域,仅写入当前批次的小数据块
完整示例代码
以下代码和你给出的ASCII写入示例功能完全一致:循环写入1~10的整数到同一个数据集,无需提前声明长度为10的全量数组。
program hdf5_append_example use hdf5 implicit none integer(hid_t) :: file_id, dset_id, dataspace_id, memspace_id integer(hsize_t) :: dims(1), max_dims(1), offset(1), count(1) integer :: i, hdferr, write_val integer, parameter :: RANK = 1 ! 初始化HDF5接口 call h5open_f(hdferr) ! 创建HDF5文件,若已存在则覆盖 call h5fcreate_f("someFile.h5", H5F_ACC_TRUNC_F, file_id, hdferr) ! 初始数据集大小为0,最大尺寸设置为无限制(支持扩展) dims = 0 max_dims = H5S_UNLIMITED_F call h5screate_simple_f(RANK, dims, dataspace_id, hdferr, max_dims) ! 创建数据集,这里用原生整数类型 call h5dcreate_f(file_id, "integer_dataset", H5T_NATIVE_INTEGER, dataspace_id, dset_id, hdferr) call h5sclose_f(dataspace_id, hdferr) ! 创建内存空间,每次仅写入1个整数,所以内存空间大小固定为1 count = 1 call h5screate_simple_f(RANK, count, memspace_id, hdferr) ! 循环写入1~10,和你给出的ASCII示例逻辑完全对应 do i = 1, 10 write_val = i ! 仅需要存储当前要写入的单个值,不需要全量数组 ! 扩展数据集:长度+1 dims = i call h5dextend_f(dset_id, dims, hdferr) ! 选中数据集里新增的最后1位位置 offset = i - 1 ! HDF5偏移量从0开始 call h5dget_space_f(dset_id, dataspace_id, hdferr) call h5sselect_hyperslab_f(dataspace_id, H5S_SELECT_SET_F, offset, count, hdferr) ! 写入当前值 call h5dwrite_f(dset_id, H5T_NATIVE_INTEGER, write_val, count, hdferr, & mem_space_id=memspace_id, file_space_id=dataspace_id) call h5sclose_f(dataspace_id, hdferr) end do ! 释放资源 call h5sclose_f(memspace_id, hdferr) call h5dclose_f(dset_id, hdferr) call h5fclose_f(file_id, hdferr) call h5close_f(hdferr) end program hdf5_append_example
编译运行说明
使用HDF5官方提供的Fortran编译包装器即可完成编译,无需手动配置链接参数:
h5fc hdf5_append_example.f90 -o hdf5_append_example ./hdf5_append_example
运行后生成的someFile.h5中integer_dataset数据集即为存储了1~10所有整数的1维数组。
批量写入优化
如果每次写入小批量数据(比如每次写100个值),仅需修改每次h5dextend_f的扩展长度和hyperslab的count参数即可,内存仅需要存储当前批次的小批量数组,依然不需要全量加载所有数据。
内容的提问来源于stack exchange,提问作者YYK
相关产品推荐
相关产品推荐

