You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Snakemake Checkpoint恢复多通配符未知输出时遇报错求助

Snakemake Checkpoint多通配符场景下的输出恢复问题解决办法

问题背景

使用Snakemake Checkpoint功能恢复预执行未知的输出文件时,当checkpoint输出路径包含{sample}这类多通配符,聚合规则的输入函数会报Missing wildcard values for sample错误,原测试代码如下:

SAMPLE = ["A", "B", "C", "D", "E"]

# a target rule to define the desired final output
rule all:
    input:
        "aggregated.txt"

# the checkpoint that shall trigger re-evaluation of the DAG
# an number of file is created in a defined directory
checkpoint somestep:
    output:
        directory("results/{sample}")
    shell:
        '''
        mkdir -p results
        mkdir results/{wildcards.sample}  # 修正原代码拼写错误mkidr→mkdir
        cd results/{wildcards.sample}
        for i in 1 2 3; do touch $i.txt; done
         '''

# input function for rule aggregate, return paths to all files produced by the checkpoint 'somestep'
def aggregate_input(wildcards):
    checkpoint_output = checkpoints.somestep.get(**wildcards).output[0]
    return expand("results/{sample}/{i}.txt", sample=wildcards.sample, i=glob_wildcards(os.path.join(checkpoint_output, "{i}.txt")).i)

rule aggregate:
    input:
        aggregate_input
    output:
        "aggregated.txt"
    shell:
        "cat {input} > {output}"

错误原因

规则aggregate本身没有定义任何通配符,因此调用aggregate_input时,wildcards对象是空的,无法获取sample值;而checkpointsomestep是针对每个sample实例化运行的,需要先收集所有checkpoint的输出结果再处理。

解决方案

修改aggregate_input函数,不依赖单个sample通配符,而是遍历所有checkpoint运行产生的输出目录,逐个解析并收集文件路径:

import os
from snakemake.io import expand, glob_wildcards

SAMPLE = ["A", "B", "C", "D", "E"]

rule all:
    input:
        "aggregated.txt"

checkpoint somestep:
    output:
        directory("results/{sample}")
    shell:
        '''
        mkdir -p results/{wildcards.sample}
        cd results/{wildcards.sample}
        for i in 1 2 3; do touch $i.txt; done
         '''

def aggregate_input(wildcards):
    # 获取所有checkpoint somestep运行产生的输出目录
    all_checkpoint_dirs = checkpoints.somestep.get_output()
    all_files = []
    for dir_path in all_checkpoint_dirs:
        # 从目录路径中解析出sample通配符值
        sample = os.path.basename(dir_path)
        # 获取当前sample目录下的所有i.txt文件
        i_vals = glob_wildcards(os.path.join(dir_path, "{i}.txt")).i
        # 生成完整文件路径并添加到列表
        all_files.extend(expand("results/{sample}/{i}.txt", sample=sample, i=i_vals))
    return all_files

rule aggregate:
    input:
        aggregate_input
    output:
        "aggregated.txt"
    shell:
        "cat {input} > {output}"

修改说明

  • 用checkpoints.somestep.get_output()替代get(**wildcards),获取所有checkpoint实例的输出目录
  • 通过os.path.basename(dir_path)从输出目录路径中提取sample值(因为目录名就是sample值)
  • 遍历每个sample的目录,收集对应i.txt文件的路径,最终返回所有文件的列表给聚合规则

验证运行

执行snakemake -j 5后,所有sample下的1.txt/2.txt/3.txt会被正确聚合到aggregated.txt中,无通配符缺失错误。

内容的提问来源于stack exchange,提问作者Ollie White

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 20:33:17