You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助编写Shell脚本批量用Prinseq处理100样本双端fastq文件

Batch Processing Paired-End FASTQ Files with Prinseq via Shell Script

Hey there! Let's solve this paired-end FASTQ batch processing problem for your 100 samples. Here's a straightforward, robust way to loop through each sample's paired files and run Prinseq automatically:

Step 1: Confirm Your File Naming Pattern

First, ensure your paired files follow a consistent naming convention. Common examples include:

  • sample001_R1.fastq / sample001_R2.fastq
  • sample001_reads1.fastq / sample001_reads2.fastq

The script below uses the _R1.fastq/_R2.fastq pattern—adjust the suffixes if your files use a different structure.

Step 2: The Shell Script

Save this as run_prinseq_batch.sh, then make it executable with chmod +x run_prinseq_batch.sh:

#!/bin/bash

# Loop through all R1 (reads1) FASTQ files
for r1 in *R1.fastq; do
    # Extract the base sample name by removing the "_R1.fastq" suffix
    sample=$(basename "$r1" _R1.fastq)
    # Define the corresponding R2 (reads2) file path
    r2="${sample}_R2.fastq"

    # Check if the R2 file exists to avoid errors
    if [ -f "$r2" ]; then
        echo "Processing sample: $sample"
        # Run Prinseq with paired files—add your custom parameters here
        perl prinseq-lite.pl -fastq "$r1" -fastq2 "$r2" \
            -min_qual_mean 20 \
            -min_len 50 \
            -out_good "${sample}_filtered"
        echo "Completed processing $sample"
    else
        echo "⚠️ Warning: R2 file $r2 not found for sample $sample. Skipping."
    fi
done

Key Details Breakdown:

  • basename "$r1" _R1.fastq: Strips the _R1.fastq suffix from the R1 filename to get your sample's core name (e.g., sample001_R1.fastq becomes sample001).
  • Double quotes around variables ("$r1", "$sample"): Ensures the script handles filenames with spaces or special characters without breaking.
  • The if [ -f "$r2" ] check: Prevents the script from crashing if a sample is missing its R2 file, and gives you a clear warning instead.
  • Customizable Prinseq parameters: Replace the example filters (-min_qual_mean, -min_len, -out_good) with whatever settings you need for your data.

Adjust for Alternative Naming Patterns

If your files use reads1/reads2 instead of R1/R2, modify these lines:

for r1 in *reads1.fastq; do
    sample=$(basename "$r1" reads1.fastq)
    r2="${sample}reads2.fastq"

Just match the suffixes to your actual file names, and the script will work seamlessly.

内容的提问来源于stack exchange,提问作者Mike F

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:02:57