You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Fortran中无需大内存数组向HDF5文件追加写入数据的实现方法

Fortran HDF5逐次追加写入数据集实现方法

核心原理

要实现类似ASCII文件逐次写入、无需全量内存数组的效果,需要使用HDF5的可扩展数据集功能:

  • 创建数据集时将维度的最大值设为H5S_UNLIMITED_F,允许后续动态修改数据集尺寸
  • 每次写入数据前先扩展数据集大小,匹配本次要写入的内容长度
  • 通过 hyperslab 选中数据集内新增的空白区域,仅写入当前批次的小数据块

完整示例代码

以下代码和你给出的ASCII写入示例功能完全一致:循环写入1~10的整数到同一个数据集,无需提前声明长度为10的全量数组。

program hdf5_append_example
  use hdf5
  implicit none
  integer(hid_t) :: file_id, dset_id, dataspace_id, memspace_id
  integer(hsize_t) :: dims(1), max_dims(1), offset(1), count(1)
  integer :: i, hdferr, write_val
  integer, parameter :: RANK = 1

  ! 初始化HDF5接口
  call h5open_f(hdferr)

  ! 创建HDF5文件,若已存在则覆盖
  call h5fcreate_f("someFile.h5", H5F_ACC_TRUNC_F, file_id, hdferr)

  ! 初始数据集大小为0,最大尺寸设置为无限制(支持扩展)
  dims = 0
  max_dims = H5S_UNLIMITED_F
  call h5screate_simple_f(RANK, dims, dataspace_id, hdferr, max_dims)

  ! 创建数据集,这里用原生整数类型
  call h5dcreate_f(file_id, "integer_dataset", H5T_NATIVE_INTEGER, dataspace_id, dset_id, hdferr)
  call h5sclose_f(dataspace_id, hdferr)

  ! 创建内存空间,每次仅写入1个整数,所以内存空间大小固定为1
  count = 1
  call h5screate_simple_f(RANK, count, memspace_id, hdferr)

  ! 循环写入1~10,和你给出的ASCII示例逻辑完全对应
  do i = 1, 10
    write_val = i ! 仅需要存储当前要写入的单个值,不需要全量数组

    ! 扩展数据集:长度+1
    dims = i
    call h5dextend_f(dset_id, dims, hdferr)

    ! 选中数据集里新增的最后1位位置
    offset = i - 1 ! HDF5偏移量从0开始
    call h5dget_space_f(dset_id, dataspace_id, hdferr)
    call h5sselect_hyperslab_f(dataspace_id, H5S_SELECT_SET_F, offset, count, hdferr)

    ! 写入当前值
    call h5dwrite_f(dset_id, H5T_NATIVE_INTEGER, write_val, count, hdferr, &
                    mem_space_id=memspace_id, file_space_id=dataspace_id)

    call h5sclose_f(dataspace_id, hdferr)
  end do

  ! 释放资源
  call h5sclose_f(memspace_id, hdferr)
  call h5dclose_f(dset_id, hdferr)
  call h5fclose_f(file_id, hdferr)
  call h5close_f(hdferr)
end program hdf5_append_example

编译运行说明

使用HDF5官方提供的Fortran编译包装器即可完成编译,无需手动配置链接参数:

h5fc hdf5_append_example.f90 -o hdf5_append_example
./hdf5_append_example

运行后生成的someFile.h5中integer_dataset数据集即为存储了1~10所有整数的1维数组。

批量写入优化

如果每次写入小批量数据(比如每次写100个值),仅需修改每次h5dextend_f的扩展长度和hyperslab的count参数即可,内存仅需要存储当前批次的小批量数组,依然不需要全量加载所有数据。

内容的提问来源于stack exchange,提问作者YYK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 09:45:00