You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake中如何用glob_wildcard处理新文件并在rule all中使用

问题:Snakemake动态处理未知数量的聚类输出文件

需求概述

构建Snakemake流程,通过cd-hit对序列聚类,拆分聚类结果文件为单个簇文件,最终为每个簇生成包含登录号的.acc文件。由于簇的数量无法预先确定,需通过通配符动态处理所有生成的簇文件。

原始Snakefile代码

#!/usr/bin/env python3

import os
import re
import pathlib
from Bio import SeqIO
from snakemake.io import glob_wildcards

#Snakemake --cores 18

###########
# GLOBALS #
###########

TRANSCRIPTOME = ["protein"]

OUTDIR = ["outdir"]

def get_cluster(wildcards):
    checkpoints.cluster_csplit.get()
    return glob_wildcards(f"{OUTDIR}/Cluster/{TRANSCRIPTOME}-{{CLUSTER}}").CLUSTER
    
#########
# RULES #
#########

rule all:
    input:
        expand('{OUTDIR}/{TRANSCRIPTOME}.90.faa.clstr', TRANSCRIPTOME = TRANSCRIPTOME, OUTDIR = OUTDIR),
        expand('{OUTDIR}/Cluster/{TRANSCRIPTOME}-0000', TRANSCRIPTOME = TRANSCRIPTOME, OUTDIR = OUTDIR),
        expand("{OUTDIR}/Cluster/{TRANSCRIPTOME}-{CLUSTER}.acc", TRANSCRIPTOME = TRANSCRIPTOME, OUTDIR = OUTDIR, CLUSTER = get_cluster)
        
########
# MAIN #
########

#Use cdhit to the cluster seqs
rule NMG_cdhit:
    input:
        '{TRANSCRIPTOME}.faa'
    output:
        fa = '{OUTDIR}/{TRANSCRIPTOME}.90.faa',
        clstr = '{OUTDIR}/{TRANSCRIPTOME}.90.faa.clstr'
    threads:
        14
    shell:
        'cd-hit '
        '-T {threads} '
        '-i {input} '
        '-o {output.fa} '
        '-d 0 '
        '-sc 1 '
        '-c 0.90 '

#split the cluster file into seperate cluster files       
checkpoint cluster_csplit:
    input:
        '{OUTDIR}/{TRANSCRIPTOME}.90.faa.clstr'
    output:
        '{OUTDIR}/Cluster/{TRANSCRIPTOME}-0000'
    params:
        '{OUTDIR}/Cluster/{TRANSCRIPTOME}-'
    shell:
        """csplit --digits=4 -z --prefix={params} {input} "/>Cluster*/" {{*}} """
 
#process       
rule get_cluster_files:
    input:
        "{OUTDIR}/Cluster/{TRANSCRIPTOME}-{CLUSTER}"
    output:
        "{OUTDIR}/Cluster/{TRANSCRIPTOME}-{CLUSTER}.acc"
    shell:
        "echo {input} > {output}" #To be replaced with script to return accession numbers. Just renames the files for now.

错误信息

MissingInputException in line 63 of /home/js/Projects/cd-hit_2_fa_snakemake/Snakefile:
Missing input files for rule get_cluster_files:
outdir/Cluster/protein-<function get_cluster at 0x7f5ca2b07b50>" 

问题分析与修复方案

核心问题

  1. 全局变量类型错误:TRANSCRIPTOME和OUTDIR被定义为列表,导致路径格式化时生成无效路径(如['outdir']/Cluster/['protein']-{CLUSTER})。
  2. Checkpoint输出定义不准确:仅指定单个拆分文件作为输出,Snakemake无法识别目录下动态生成的所有簇文件。
  3. expand函数中传递通配符的方式错误:直接传入get_cluster函数,但未正确传递wildcards参数,导致函数未被执行,反而被当作字符串拼接。

修改后的完整代码

#!/usr/bin/env python3

import os
import re
import pathlib
from Bio import SeqIO
from snakemake.io import glob_wildcards

#Snakemake --cores 18

###########
# GLOBALS #
###########

# 改为字符串而非列表,避免路径格式化错误
TRANSCRIPTOME = "protein"
OUTDIR = "outdir"

def get_cluster(wildcards):
    # 等待checkpoint完成,确保所有簇文件已生成
    checkpoints.cluster_csplit.get()
    # 正确格式化路径,获取所有簇编号
    clusters = glob_wildcards(f"{OUTDIR}/Cluster/{TRANSCRIPTOME}-{{CLUSTER}}").CLUSTER
    return clusters
    
#########
# RULES #
#########

rule all:
    input:
        f"{OUTDIR}/{TRANSCRIPTOME}.90.faa.clstr",
        # 不再单独指定单个簇文件,而是通过get_cluster获取所有动态生成的目标文件
        expand("{OUTDIR}/Cluster/{TRANSCRIPTOME}-{CLUSTER}.acc", CLUSTER=get_cluster)
        
########
# MAIN #
########

#Use cdhit to cluster seqs
rule NMG_cdhit:
    input:
        f"{TRANSCRIPTOME}.faa"
    output:
        fa = f"{OUTDIR}/{TRANSCRIPTOME}.90.faa",
        clstr = f"{OUTDIR}/{TRANSCRIPTOME}.90.faa.clstr"
    threads:
        14
    shell:
        'cd-hit '
        '-T {threads} '
        '-i {input} '
        '-o {output.fa} '
        '-d 0 '
        '-sc 1 '
        '-c 0.90 '

#split the cluster file into separate cluster files       
checkpoint cluster_csplit:
    input:
        f"{OUTDIR}/{TRANSCRIPTOME}.90.faa.clstr"
    output:
        # 指定输出为整个Cluster目录,告知Snakemake该目录下有动态生成的文件
        directory(f"{OUTDIR}/Cluster")
    params:
        prefix = f"{OUTDIR}/Cluster/{TRANSCRIPTOME}-"
    shell:
        # 先创建Cluster目录,避免csplit报错
        "mkdir -p {output}; "
        "csplit --digits=4 -z --prefix={params.prefix} {input} \"/>Cluster*/\" {{*}} "
 
#process each cluster file to generate .acc file       
rule get_cluster_files:
    input:
        f"{OUTDIR}/Cluster/{TRANSCRIPTOME}-{{CLUSTER}}"
    output:
        f"{OUTDIR}/Cluster/{TRANSCRIPTOME}-{{CLUSTER}}.acc"
    shell:
        "echo {input} > {output}" # 替换为提取登录号的脚本

关键修改说明

  1. 全局变量调整:将TRANSCRIPTOME和OUTDIR改为字符串,简化路径处理,若后续需处理多个样本,可改为列表并结合expand循环处理。
  2. Checkpoint输出改为目录:用directory()指定Cluster目录作为输出,确保Snakemake跟踪该目录下所有动态生成的文件。
  3. rule all简化:移除对单个簇文件的硬编码依赖,完全通过get_cluster函数动态获取所有需要生成的.acc文件。
  4. 新增目录创建:在checkpoint的shell命令中添加mkdir -p {output},避免csplit因目录不存在报错。
  5. 路径格式化优化:使用f-string直接生成固定路径,减少expand的不必要使用,提高可读性。

内容的提问来源于stack exchange,提问作者J-Skelly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 11:10:00