求助编写Shell脚本批量用Prinseq处理100样本双端fastq文件
Hey there! Let's solve this paired-end FASTQ batch processing problem for your 100 samples. Here's a straightforward, robust way to loop through each sample's paired files and run Prinseq automatically:
Step 1: Confirm Your File Naming Pattern
First, ensure your paired files follow a consistent naming convention. Common examples include:
sample001_R1.fastq/sample001_R2.fastqsample001_reads1.fastq/sample001_reads2.fastq
The script below uses the _R1.fastq/_R2.fastq pattern—adjust the suffixes if your files use a different structure.
Step 2: The Shell Script
Save this as run_prinseq_batch.sh, then make it executable with chmod +x run_prinseq_batch.sh:
#!/bin/bash # Loop through all R1 (reads1) FASTQ files for r1 in *R1.fastq; do # Extract the base sample name by removing the "_R1.fastq" suffix sample=$(basename "$r1" _R1.fastq) # Define the corresponding R2 (reads2) file path r2="${sample}_R2.fastq" # Check if the R2 file exists to avoid errors if [ -f "$r2" ]; then echo "Processing sample: $sample" # Run Prinseq with paired files—add your custom parameters here perl prinseq-lite.pl -fastq "$r1" -fastq2 "$r2" \ -min_qual_mean 20 \ -min_len 50 \ -out_good "${sample}_filtered" echo "Completed processing $sample" else echo "⚠️ Warning: R2 file $r2 not found for sample $sample. Skipping." fi done
Key Details Breakdown:
basename "$r1" _R1.fastq: Strips the_R1.fastqsuffix from the R1 filename to get your sample's core name (e.g.,sample001_R1.fastqbecomessample001).- Double quotes around variables (
"$r1","$sample"): Ensures the script handles filenames with spaces or special characters without breaking. - The
if [ -f "$r2" ]check: Prevents the script from crashing if a sample is missing its R2 file, and gives you a clear warning instead. - Customizable Prinseq parameters: Replace the example filters (
-min_qual_mean,-min_len,-out_good) with whatever settings you need for your data.
Adjust for Alternative Naming Patterns
If your files use reads1/reads2 instead of R1/R2, modify these lines:
for r1 in *reads1.fastq; do sample=$(basename "$r1" reads1.fastq) r2="${sample}reads2.fastq"
Just match the suffixes to your actual file names, and the script will work seamlessly.
内容的提问来源于stack exchange,提问作者Mike F

