You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Snakemake为带通配符的Checkpoint生成单个汇总文件?

解决Snakemake带通配符Checkpoint生成单个汇总文件的问题

问题背景

原本的工作流中,aggregate规则会为每个样本单独生成汇总文件,现在需要生成单个汇总文件/目录,但直接修改aggregate输出为非通配符路径时,出现错误:

Error:
  WorkflowError:
    Missing wildcard values for sample
Wildcards:

原因是checkpoints.clustering.get(**wildcards)需要sample通配符参数,但无通配符的aggregate规则无法提供该参数,导致绑定失败。

解决方案

方法一:直接收集所有样本的Checkpoint输出并汇总

修改输入函数,主动获取所有样本列表,遍历每个样本的Checkpoint输出,收集所有中间文件后汇总:

import os
from snakemake.io import glob_wildcards, expand

# 目标规则:指定最终输出为单个文件
rule all:
    input:
        "aggregated.txt"

# 带通配符的Checkpoint,按样本生成聚类结果
checkpoint clustering:
    input:
        "samples/{sample}.txt"
    output:
        clusters=directory("clustering/{sample}")
    shell:
        "mkdir -p clustering/{wildcards.sample}; "
        "for i in 1 2 3; do echo $i > clustering/{wildcards.sample}/$i.txt; done"

# 中间处理规则:转换聚类文件路径
rule intermediate:
    input:
        "clustering/{sample}/{i}.txt"
    output:
        "post/{sample}/{i}.txt"
    shell:
        "cp {input} {output}"

# 修改输入函数:主动获取所有样本,收集所有中间文件
def aggregate_input(wildcards):
    # 获取samples目录下的所有样本
    samples = glob_wildcards("samples/{sample}.txt").sample
    all_post_files = []
    
    for sample in samples:
        # 获取当前样本的Checkpoint输出目录
        checkpoint_dir = checkpoints.clustering.get(sample=sample).output[0]
        # 收集该样本下的所有post文件
        cluster_ids = glob_wildcards(os.path.join(checkpoint_dir, "{i}.txt")).i
        sample_files = expand("post/{sample}/{i}.txt", sample=sample, i=cluster_ids)
        all_post_files.extend(sample_files)
    
    return all_post_files

# 单个汇总规则:合并所有中间文件
rule aggregate:
    input:
        aggregate_input
    output:
        "aggregated.txt"
    shell:
        "cat {input} > {output}"

方法二:先按样本汇总,再合并所有样本结果

保留原有的按样本汇总逻辑,新增一个顶层规则合并所有样本的汇总文件,适合需要保留单个样本汇总结果的场景:

import os
from snakemake.io import glob_wildcards, expand

# 目标规则:指定最终合并后的文件
rule all:
    input:
        "merged_aggregated.txt"

# 带通配符的Checkpoint,按样本生成聚类结果
checkpoint clustering:
    input:
        "samples/{sample}.txt"
    output:
        clusters=directory("clustering/{sample}")
    shell:
        "mkdir -p clustering/{wildcards.sample}; "
        "for i in 1 2 3; do echo $i > clustering/{wildcards.sample}/$i.txt; done"

# 中间处理规则:转换聚类文件路径
rule intermediate:
    input:
        "clustering/{sample}/{i}.txt"
    output:
        "post/{sample}/{i}.txt"
    shell:
        "cp {input} {output}"

# 原输入函数:按样本收集中间文件
def aggregate_input(wildcards):
    checkpoint_output = checkpoints.clustering.get(**wildcards).output[0]
    return expand("post/{sample}/{i}.txt",
           sample=wildcards.sample,
           i=glob_wildcards(os.path.join(checkpoint_output, "{i}.txt")).i)

# 按样本汇总规则:保留原逻辑
rule aggregate:
    input:
        aggregate_input
    output:
        "aggregated/{sample}.txt"
    shell:
        "cat {input} > {output}"

# 新增合并规则:合并所有样本的汇总文件
rule merge_aggregates:
    input:
        expand("aggregated/{sample}.txt", sample=glob_wildcards("samples/{sample}.txt").sample)
    output:
        "merged_aggregated.txt"
    shell:
        "cat {input} > {output}"

关键思路

  • 核心问题是无通配符的规则无法为checkpoints.clustering.get()提供sample参数,因此需要主动获取所有样本列表,或拆分汇总步骤避免直接绑定无通配符规则与带通配符Checkpoint。
  • 使用glob_wildcards("samples/{sample}.txt")预先获取样本列表时,需确保samples目录下的样本文件在工作流启动时已存在。

内容的提问来源于stack exchange,提问作者Yannis Schöneberg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 14:04:55