基于中间相似ID合并fastq.gz文件的Bash脚本实现疑问
Hey there! Let's sort out this file grouping and merging task for your fastq.gz files. Here's how you can complete that missing XXXX section in your Bash script to group files by those HXX IDs (like H10, H11) and merge them correctly:
Complete Bash Script with ID Grouping
# Initialize an associative array to map each HXX ID to its list of files declare -A ids # Loop through all your target fastq.gz files (matching the 01_002_ prefix) for file in 01_002_*.fastq.gz; do # Extract the HXX ID from the filename # Step 1: Remove the fixed prefix "01_002_" temp="${file#01_002_}" # Step 2: Grab everything before the first underscore (gives us H10/H11 etc.) id="${temp%%_*}" # Append the current file to the list for this ID ids["$id"]+="$file " done # Now iterate over each ID and merge its corresponding files for id in "${!ids[@]}"; do # Define the output filename for merged files merged_output="01_002_${id}_merged.fastq.gz" # Merge the files directly with cat (works for gzipped files since gzip is stream-based) echo "Merging files for ID *$id* into *$merged_output*..." cat ${ids["$id"]} > "$merged_output" # Optional: Verify the merged gzip file isn't corrupted echo "Checking integrity of $merged_output..." gunzip -t "$merged_output" done
Key Details Explained
- ID Extraction: We use Bash parameter expansion to pull out the
HXXID without relying on external tools likesedorawk:${file#01_002_}strips off the fixed prefix from the filename${temp%%_*}trims everything from the first underscore onwards, leaving just theH10/H11ID
- Associative Array: The
idsarray stores all files belonging to each ID, so we can easily loop through groups later - Merging Gzipped Files: Using
catdirectly on.fastq.gzfiles works perfectly because gzip uses stream compression—no need to decompress first, which saves time and disk space - Integrity Check: The optional
gunzip -tcommand verifies that the merged file is a valid gzip archive, so you can catch any issues right away
This script will automatically group all your files by their HXX IDs and merge each pair into a single, valid gzipped file.
内容的提问来源于stack exchange,提问作者user2300940
相关产品推荐
相关产品推荐

