You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake中如何实现单输入多输出的‘发散型’规则?

Snakemake中“发散型”规则的最优实现方案

收敛型规则定义与示例

我将“收敛型”规则定义为从多个输入生成单个输出的规则,示例代码如下:

group2samples = {
    "A": ["s1", "s2"],
    "B": ["s3", "s4"]}

rule all:
    input: [f"{group}.txt" for group in group2samples]

def set_input(wildcards):
    return [f"{sample}.txt" for sample in group2samples[wildcards.group]]

rule converging:
    input:
        set_input
    output:
        "{group}.txt"
    shell:
        "cat {input} > {output}"

发散型规则的需求与无效尝试

现在需要实现从单个输入生成多个输出的“发散型”规则,我尝试了以下写法,但无法生效:

group2samples = {
    "A": ["s1", "s2"],
    "B": ["s3", "s4"]}

rule all:
    input:
        [
            f"{group}/{sample}.txt"
            for sample in group2samples[group]
            for group in group2samples]

rule diverging:
    input:
        "{group}.txt"
    output:
        # 尝试用lambda动态生成输出列表,但Snakemake不支持这种写法
        lambda wildcards: [f"{{group}}/{sample}.txt" for sample in group2samples[wildcards.group]]
    shell:
        "my_data_extracting_script.py {input}"

现有不理想的方案

我想到两种可行但存在缺陷的实现方式:

  • 将输出文件打包为归档文件作为规则输出,但无法得到独立的单个文件
  • 使用目录作为输出,但不符合显式文件列表的需求,且Snakemake文档不推荐使用directory参数

更优实现方式:使用Checkpoint

Snakemake官方推荐使用checkpoint处理这类“单输入多输出”的场景,既能生成独立文件,又满足显式文件列表的要求。具体实现如下:

group2samples = {
    "A": ["s1", "s2"],
    "B": ["s3", "s4"]}

# 定义最终需要生成的所有文件
rule all:
    input:
        expand("{group}/{sample}.txt", group=group2samples.keys(), sample=group2samples)

# 用checkpoint执行生成多输出的脚本,先将文件输出到临时目录
checkpoint diverging:
    input:
        "{group}.txt"
    output:
        directory("{group}_temp")
    shell:
        """
        mkdir -p {output}
        # 假设你的脚本支持指定输出目录,把生成的文件放到临时目录里
        my_data_extracting_script.py {input} --output-dir {output}
        """

# 定义收集输出文件的函数,根据group获取对应的sample列表
def collect_output_files(wildcards):
    samples = group2samples[wildcards.group]
    return expand("{group}/{sample}.txt", group=wildcards.group, sample=samples)

# 将临时目录中的文件移动到最终路径,同时让Snakemake追踪每个独立文件
rule move_to_final:
    input:
        checkpoint(diverging).output,
        files=collect_output_files
    output:
        "{group}/{sample}.txt"
    shell:
        "mv {input[0]}/{wildcards.sample}.txt {output}"

方案说明

  1. Checkpoint:负责执行生成多输出的脚本,将临时输出存放在指定目录中,Snakemake会追踪这个目录的更新
  2. 收集函数:根据通配符动态获取每个group对应的所有输出文件路径
  3. 移动规则:将临时目录中的文件转移到最终目标路径,确保每个文件都被Snakemake显式追踪,满足需求

内容的提问来源于stack exchange,提问作者bli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 11:43:15