如何对结构相同的两个H5文件执行元素级相加并保存到新文件
Got it, let's sort out this problem for you! The error you're seeing makes total sense—h5py's File objects aren't arrays, so you can't just add them directly with +. Instead, you need to traverse each dataset inside the files and perform element-wise addition on the actual data stored in those datasets.
Here's a step-by-step solution that works for structure-matching H5 files (like your 1.h5 and 2.h5):
Full Working Code
import h5py import numpy as np # Use context managers to handle file opening/closing automatically with h5py.File('1.h5', 'r') as file1, h5py.File('2.h5', 'r') as file2, h5py.File('sum_result.h5', 'w') as output_file: # Define a function to process each object in the H5 file def process_dataset(name, obj): if isinstance(obj, h5py.Dataset): # Read the full dataset from both input files data1 = file1[name][()] data2 = file2[name][()] # Perform element-wise addition summed_data = data1 + data2 # Create the dataset in the output file, preserving original dtype/shape output_dset = output_file.create_dataset(name, data=summed_data, dtype=obj.dtype, shape=obj.shape) # Copy over all metadata attributes from the original dataset for attr_name, attr_value in obj.attrs.items(): output_dset.attrs[attr_name] = attr_value elif isinstance(obj, h5py.Group): # If it's a group (folder-like structure), create the same group in the output output_file.create_group(name) # Traverse all items in the first file and process them file1.visititems(process_dataset)
Key Details Explained
- Context Managers: Using
withensures files are closed properly even if something goes wrong. - Dataset vs Group Check: The
visititemsmethod iterates over every object in the file—we only want to add data inDatasetobjects, while we just replicateGroupstructures to keep the output file's hierarchy the same. - Preserve Metadata: We copy over all attributes from the original datasets (like dtype, shape, or custom metadata) so the output file matches the input structure as closely as possible.
For Large Datasets (Memory-Friendly)
If your H5 files are too big to load entirely into memory, use chunked reading/writing to process data in smaller pieces:
import h5py import numpy as np with h5py.File('1.h5', 'r') as file1, h5py.File('2.h5', 'r') as file2, h5py.File('sum_result.h5', 'w') as output_file: def process_large_dataset(name, obj): if isinstance(obj, h5py.Dataset): # Create output dataset with the same chunking as the original output_dset = output_file.create_dataset( name, shape=obj.shape, dtype=obj.dtype, chunks=obj.chunks ) # Iterate over each chunk and add data piece by piece for chunk_slice in obj.iter_chunks(): data1 = file1[name][chunk_slice] data2 = file2[name][chunk_slice] output_dset[chunk_slice] = data1 + data2 # Copy metadata attributes for attr_name, attr_value in obj.attrs.items(): output_dset.attrs[attr_name] = attr_value elif isinstance(obj, h5py.Group): output_file.create_group(name) file1.visititems(process_large_dataset)
This code will correctly compute the element-wise sum of every matching dataset in your two H5 files and save the result to sum_result.h5.
内容的提问来源于stack exchange,提问作者Hitesh

