编程新手求助:求基于h5py的Python脚本遍历目录合并H5文件
Hey there! As someone who’s worked with HDF5 files quite a bit, I’ve put together a beginner-friendly Python script using h5py that does exactly what you need: it crawls through every directory and subdirectory under a target path, collects all H5 files, and merges them into a single output file. Let’s dive in!
The Complete Script
import os import h5py def merge_h5_files(root_dir, output_filename): # Create the output H5 file in write mode with h5py.File(output_filename, 'w') as output_h5: # Walk through every folder and subfolder in the target directory for dirpath, _, filenames in os.walk(root_dir): for filename in filenames: # Only process files with .h5 or .hdf5 extensions (case-insensitive) if filename.lower().endswith(('.h5', '.hdf5')): file_path = os.path.join(dirpath, filename) print(f"Processing file: {file_path}") try: # Open the input H5 file in read-only mode with h5py.File(file_path, 'r') as input_h5: # Helper function to copy groups and datasets def copy_items(name, obj): # Handle datasets if isinstance(obj, h5py.Dataset): # Add filename prefix to avoid name conflicts new_name = f"{os.path.splitext(filename)[0]}/{name}" # Copy the dataset to output, preserving its properties input_h5.copy(obj, output_h5, name=new_name) # Handle groups (folders inside the H5 file) elif isinstance(obj, h5py.Group): # Create the group in output if it doesn't exist new_group_name = f"{os.path.splitext(filename)[0]}/{name}" if new_group_name not in output_h5: output_h5.create_group(new_group_name) # Traverse every item in the input file and copy it input_h5.visititems(copy_items) except Exception as e: # Skip problematic files instead of crashing the whole script print(f"Error processing {file_path}: {str(e)}") continue if __name__ == "__main__": # Replace this with the path to your folder of H5 files target_directory = "./your_h5_files_folder" # Name your merged output file output_file = "merged_all.h5" merge_h5_files(target_directory, output_file) print(f"Merging done! Your combined file is at {output_file}")
How This Script Works
Let’s break down the key parts so you understand what’s happening:
- Directory Crawling:
os.walk()goes through every folder and subfolder under your target directory, so no H5 file gets missed. - File Filtering: We only process files ending with
.h5or.hdf5(works for uppercase extensions like.H5too). - Conflict Prevention: If two files have datasets/groups with the same name (e.g., both have a
datadataset), we add the source filename as a prefix (sofile1.h5’sdatabecomesfile1/datain the merged file). This keeps everything organized. - Error Handling: If a file is corrupted or you don’t have permission to read it, the script prints an error and moves on instead of crashing.
- Structure Preservation: The
visititems()method lets us copy every group and dataset from the input files, keeping their original hierarchy intact.
Important Notes for Beginners
- Backup First: Always make copies of your original H5 files before merging! Merging is irreversible, and you don’t want to lose data if something goes wrong.
- Memory Check: If you’re merging huge H5 files, keep an eye on your system’s memory. This script copies datasets temporarily—for extra-large files, you might want to look into chunked copying, but this works for most standard use cases.
- Customize as Needed: If you’re 100% sure all dataset/group names are unique across files, you can remove the
{os.path.splitext(filename)[0]}/part fromnew_nameandnew_group_nameto skip the filename prefix. - Path Setup: Don’t forget to replace
./your_h5_files_folderwith your actual folder path. On Windows, use paths likeC:/Users/YourName/H5Files(forward slashes work) or double backslashes (C:\\Users\\YourName\\H5Files).
内容的提问来源于stack exchange,提问作者Sameer
相关产品推荐
相关产品推荐

