如何在Snakemake Shell的单引号与占位符中使用变量
Snakemake Shell占位符与Bash变量冲突的解决办法
问题重现
以下是简化后的Snakemake规则:
rule test: output: output_tex = expand("TPA_{region}.tex", region = [0,1]), shell: """ for((i=0;i<2;i++)); do echo '"{output.output_tex['$i']}"'; done """
预期输出:
"TPA_0.tex" "TPA_1.tex"
问题核心:在Snakemake的shell脚本中,尝试用'$i'引用Bash循环变量来索引Snakemake的输出列表时,会因为语法歧义导致解析失败,无法得到预期结果。
可行解决方案
方案1:将Snakemake输出转为Bash数组
先把Snakemake生成的输出列表赋值给Bash数组,再在循环中通过数组索引调用:
rule test: output: output_tex = expand("TPA_{region}.tex", region = [0,1]), shell: """ FILES=({output.output_tex}) for ((i=0; i<2; i++)); do echo "\"${FILES[$i]}\"" done """
原理:({output.output_tex})会将Snakemake展开的多个文件名直接转为Bash数组元素,后续用标准的Bash数组索引${FILES[$i]}即可正常获取对应文件路径。
方案2:利用Snakemake内置循环特性
如果不需要复杂的Bash循环逻辑,直接使用Snakemake对多输出的自动循环展开:
rule test: output: output_tex = expand("TPA_{region}.tex", region = [0,1]), shell: """ echo '"{output}"' """
原理:当{output}对应多个输出文件时,Snakemake会自动遍历每个文件路径并执行shell命令,直接输出预期结果。
方案3:通过eval解析变量(慎用)
如果必须保留Bash循环的复杂逻辑,可借助eval强制解析顺序:
rule test: output: output_tex = expand("TPA_{region}.tex", region = [0,1]), shell: """ for i in {0..1}; do eval 'echo "\"{output.output_tex[$i]}\""' done """
注意:eval存在安全风险,若输出文件名包含特殊字符可能引发问题,优先推荐前两种方案。
内容的提问来源于stack exchange,提问作者RedQuark
相关产品推荐
相关产品推荐

