如何在Bash脚本中根据samplesheet的Group列执行对应函数
问题:根据samplesheet的Group列匹配执行对应Bash函数
我有一个包含三个函数的Bash脚本,需要根据samplesheet.txt中Group列的信息决定执行哪个函数:Group=Both时执行process_both,Group=DNA时执行process_dna,Group=RNA时执行process_rna。当前脚本的执行逻辑未按Group列匹配,请问如何修改脚本实现该功能?
附samplesheet.txt内容
DNA RNA purity Group sample_DNA_1 sample_RNA_1 0.8 Both sample_DNA_2 sample_RNA_2 0.5 Both NA sample_RNA_4 0.3 RNA sample_DNA_3 NA 0.1 DNA
现有代码
process_both() { local dna=$1 local rna=$2 local purity=$3 singularity exec \ sing.sif \ ... ... dna_id {dna} \ rna_id {rna} \ purity {purity} \ ... } process_dna() { local dna=$1 local rna=$2 local purity=$3 singularity exec \ sing.sif \ ... ... dna_id {dna} \ purity {purity} \ ... } process_rna() { local dna=$1 local rna=$2 local purity=$3 singularity exec \ sing.sif \ ... ... rna_id {rna} \ purity {purity} \ ... } # Run the functions tail +2 samplesheet.txt | while read dna rna purity do process_both "$dna" "$rna" "$purity" done tail +2 samplesheet.txt | while read dna rna purity do process_rna "$dna" "$rna" "$purity" done tail +2 samplesheet.txt | while read dna rna purity do process_dna "$dna" "$rna" "$purity" done
解决方案
原脚本的问题在于:三次重复读取samplesheet,且完全没有依据Group列做判断,导致所有函数会被全量执行。修改思路是单次循环读取所有字段,根据Group值分支调用对应函数,修改后的完整代码如下:
process_both() { local dna=$1 local rna=$2 local purity=$3 singularity exec \ sing.sif \ ... ... dna_id ${dna} \ rna_id ${rna} \ purity ${purity} \ ... } process_dna() { local dna=$1 local rna=$2 local purity=$3 singularity exec \ sing.sif \ ... ... dna_id ${dna} \ purity ${purity} \ ... } process_rna() { local dna=$1 local rna=$2 local purity=$3 singularity exec \ sing.sif \ ... ... rna_id ${rna} \ purity ${purity} \ ... } # Run the functions tail +2 samplesheet.txt | while read dna rna purity group do case "${group}" in "Both") process_both "$dna" "$rna" "$purity" ;; "DNA") process_dna "$dna" "$rna" "$purity" ;; "RNA") process_rna "$dna" "$rna" "$purity" ;; *) echo "Unknown Group value: ${group} for DNA=${dna}, RNA=${rna}" >&2 ;; esac done
关键修改点说明
- 读取Group字段:在
read命令中添加group变量,确保读取到每行的Group列值 - 单次循环处理:只读取一次samplesheet,避免重复IO操作
- 分支判断逻辑:用
case语句根据Group值精准调用对应函数,同时增加默认分支处理未知Group值的情况(可选,用于错误排查) - 修正变量引用:原函数内的
{dna}需改为${dna}才能正确解析Bash变量(原代码存在语法错误)
内容的提问来源于stack exchange,提问作者user2300940
相关产品推荐
相关产品推荐

