如何批量合并同个体对应的R1与R2文件?
Absolutely! Batch processing is the perfect solution here—you’ll cut out all that tedious manual work in no time. Let’s break this down into two practical scenarios based on your setup, since you mentioned having unique sample IDs and a related file (I’m assuming that’s a list of your sample IDs; adjust accordingly if it’s formatted differently).
Scenario 1: Merge using consistent filenames
If your files follow a predictable naming pattern like {SampleID}_R1.fastq and {SampleID}_R2.fastq (replace .fastq with your actual file extension, e.g., .fq.gz), you can use a simple bash loop to extract unique IDs and merge pairs automatically:
For uncompressed files (e.g., .fastq, .txt)
# Extract unique sample IDs from R1 filenames for sample_id in $(ls *R1.fastq | sed 's/_R1\.fastq$//'); do # Merge R1 and R2 into a single file for the sample cat "${sample_id}_R1.fastq" "${sample_id}_R2.fastq" > "${sample_id}_merged.fastq" done
For compressed files (e.g., .fastq.gz, .fq.gz)
If your files are gzipped, use zcat to avoid decompressing them first (saves disk space and processing time):
for sample_id in $(ls *R1.fastq.gz | sed 's/_R1\.fastq\.gz$//'); do zcat "${sample_id}_R1.fastq.gz" "${sample_id}_R2.fastq.gz" | gzip > "${sample_id}_merged.fastq.gz" done
Scenario 2: Merge using a sample ID list file
If you have a dedicated file (like sample_ids.txt) with one unique ID per line, this method is more reliable—it avoids issues with unexpected filename patterns or stray files in your folder:
With error checking (recommended)
This version skips samples where either R1 or R2 is missing and prints clear status messages:
while read -r sample_id; do # Define paths to R1 and R2 files r1_file="${sample_id}_R1.fastq" r2_file="${sample_id}_R2.fastq" # Only merge if both files exist if [[ -f "$r1_file" && -f "$r2_file" ]]; then cat "$r1_file" "$r2_file" > "${sample_id}_merged.fastq" echo "Successfully merged: $sample_id" else echo "Warning: Skipping $sample_id - missing $r1_file or $r2_file" fi done < sample_ids.txt
Quick tips to avoid mistakes
- Test first: Run the loop with 1-2 sample IDs first (e.g., modify the loop to only process
ID001andID002) to verify merged files look correct. - Backup originals: Make sure you have copies of your R1/R2 files before batch merging—better safe than sorry!
- Adjust extensions: Swap
.fastqor.fastq.gzwith your actual file type (e.g.,.txt,.fq) wherever needed.
内容的提问来源于stack exchange,提问作者Anna1364

