Snakemake中为每个通配符正确获取CSV行数参数的问题
解决Snakemake中为每个通配符文件获取行数并传入参数的问题
问题场景
在Snakemake工作流中,有一个R脚本scripts/use_lines.R需要接收多个带表头CSV文件(a.csv、b.csv、c.csv)的条目数作为参数。由于每个CSV的每行都包含字符串something,尝试通过规则的shell命令获取该数值时,通配符解析出错,生成的命令变为grep -c something ../resources/a b c.csv,而非针对每个通配符单独生成grep -c something ../resources/a.csv这类命令。
解决方案
方法1:通过Python表达式直接计算行数(推荐)
利用Snakemake的params字段结合Python代码,直接在工作流中计算每个文件匹配something的行数,无需调用shell命令,彻底避免通配符解析问题:
rule run_r_script: input: csv="../resources/{sample}.csv" output: "output/{sample}.txt" params: # 遍历文件每行,统计包含"something"的行数 n_lines=lambda wildcards: sum(1 for line in open(f"../resources/{wildcards.sample}.csv") if "something" in line) shell: """ Rscript scripts/use_lines.R --n_lines {params.n_lines} --input {input.csv} --output {output} """
方法2:在shell中针对单个文件调用grep
如果必须使用grep,确保通过{input.csv}传递完整文件路径,让Snakemake自动替换通配符为正确的文件名:
rule run_r_script: input: csv="../resources/{sample}.csv" output: "output/{sample}.txt" shell: """ # 针对单个文件统计行数 N_LINES=$(grep -c "something" {input.csv}) Rscript scripts/use_lines.R --n_lines $N_LINES --input {input.csv} --output {output} """
方法3:批量处理多个文件(循环遍历)
若需要一次处理所有CSV文件,可通过shell循环逐个计算每个文件的行数,再将所有行数作为参数传入R脚本:
rule run_r_batch: input: # 生成所有CSV文件路径 csvs=expand("../resources/{sample}.csv", sample=["a", "b", "c"]) output: "output/batch_results.txt" shell: """ N_LINES_LIST="" # 遍历每个输入文件,统计行数并收集结果 for FILE in {input.csvs}; do COUNT=$(grep -c "something" $FILE) N_LINES_LIST="$N_LINES_LIST $COUNT" done # 将所有行数传入R脚本 Rscript scripts/use_lines.R --n_lines_list "$N_LINES_LIST" --input_files {input.csvs} --output {output} """
内容的提问来源于stack exchange,提问作者Giulio Centorame
相关产品推荐
相关产品推荐

