如何在Bash脚本中用for循环批量处理配对Fastq文件并重命名
批量处理配对末端Fastq文件的Bash脚本优化
需求说明
需要批量处理目录中17组配对末端Fastq文件(共34个),核心需求:
- 脚本循环时自动替换输入输出文件名的前缀(如处理
file_002时,所有输出文件前缀变为file_002) - 仅配对合并同编号的R1和R2文件(如
file_001_R1.fastq.gz与file_001_R2.fastq.gz配对处理)
原脚本问题
原脚本中所有文件名写死为file_001,无法批量处理不同编号的样本,循环逻辑也未正确遍历目标文件。
优化后的脚本
# 遍历目录下所有R1文件,提取样本前缀 for r1_file in directory_name/*_R1.fastq.gz; do # 提取样本前缀(如从file_001_R1.fastq.gz得到file_001) sample=$(basename "$r1_file" _R1.fastq.gz) # 对应的R2文件路径 r2_file="directory_name/${sample}_R2.fastq.gz" # 跳过不存在对应R2的样本(可选,确保配对完整性) if [ ! -f "$r2_file" ]; then echo "警告:未找到${sample}对应的R2文件,跳过" continue fi # 执行PEAR合并配对文件 pear -f "$r1_file" -r "$r2_file" -o "${sample}.fastq" # 提取Barcode序列 cutadapt -g TGATAACAATTGGAGCAGCCTC...GGATCGACCAAGAACCAGCA -o "${sample}_barcode.fastq" "${sample}.fastq" # 提取UMI序列 cutadapt -g GTGTACAAATAATTGTCAAC...CTGTCTCTTATACACATCTC -o "${sample}_UMI.fastq" "${sample}.fastq" # 合并Barcode和UMI文件 seqkit concat "${sample}_barcode.fastq" "${sample}_UMI.fastq" > "${sample}_concatenation.fastq" # 去除重复序列 seqkit rmdup -s "${sample}_concatenation.fastq" -o "${sample}_unique_pairs.fastq" # 提取子序列生成fasta文件 seqkit subseq -r "${sample}_unique_pairs.fastq" > "${sample}_unique_barcodes.fasta" # Bowtie比对 bowtie -q --suppress 1,2,4,6,7,8 -x ref_index "${sample}_unique_barcodes.fasta" > "${sample}_barcodes_allignment.bowtie" # 统计比对结果 sort "${sample}_barcodes_allignment.bowtie" | uniq -c > "${sample}_barcode_counts.txt" # 转换为CSV格式 awk 'BEGIN{print "Barcode,TF_variant,Code"}{print $3","$2","$1}' "${sample}_barcode_counts.txt" > "${sample}_barcode_counts.csv" done
关键逻辑说明
- 动态获取样本前缀:通过
basename "$r1_file" _R1.fastq.gz提取R1文件名的前缀,确保每个循环处理唯一的样本编号 - 配对文件校验:通过判断
$r2_file是否存在,避免处理缺失配对的样本 - 变量统一替换:所有输入输出文件名使用
${sample}变量,确保每个样本的处理流程生成对应前缀的文件 - 遍历逻辑修正:原脚本遍历
directory_name目录本身,改为遍历目录下的所有R1文件,确保只处理配对样本
内容的提问来源于stack exchange,提问作者Jinsoul Jung
相关产品推荐
相关产品推荐

