You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake规则迭代执行问题:如何逐行处理实验样本数据

解决Snakemake批量处理样本时的任务合并问题

嘿,我懂你这种挫败感——本来以为摸透了Snakemake,结果处理多样本的时候直接翻车,所有参数被合并成一个任务跑了。其实问题出在你把整个DataFrame的列直接传给了规则的input和output,Snakemake会把这当成一个大任务,而不是逐个样本执行。下面给你改好的方案,一步一步帮你搞定:

修改后的完整Snakefile

import pandas as pd
configfile: "config.json"

# 读取样本表,整理每个样本的关键信息
experiments = pd.read_csv(config["experiments"], sep='\t').fillna('')
# 生成唯一的样本名(沿用你原来的提取逻辑,确保每个样本名不重复)
experiments['Name'] = [filename.split('/')[-1].split('.fa')[0].split('_R')[0] for filename in experiments['Files'].str.split(',').str[0]]
# 建立样本名到输入reads的映射字典
sample_reads_map = dict(zip(experiments['Name'], experiments['Files'].str.split(',')))
# 建立样本名到数据类型的映射字典
sample_datatype_map = dict(zip(experiments['Name'], experiments['Data type']))

rule all:
    input:
        # 为每个样本生成对应的双端输出文件
        expand(
            "{output}/Preprocess/Trimmomatic/quality_trimmed_{name}{fr}.fq",
            output=config["output"],
            name=experiments['Name'],
            fr=['_forward_paired', '_reverse_paired']
        )

rule preprocess:
    # 用{name}通配符区分不同样本,同时限制只能是表中存在的样本名
    wildcard_constraints:
        name="|".join(experiments['Name'])
    input:
        # 根据当前任务的样本名,获取对应的输入reads
        lambda wildcards: sample_reads_map[wildcards.name]
    output:
        # 为当前样本生成专属的输出文件路径
        expand(
            "{output}/Preprocess/Trimmomatic/quality_trimmed_{name}{fr}.fq",
            output=config["output"],
            name=wildcards.name,
            fr=['_forward_paired', '_reverse_paired']
        )
    threads: config["threads"]
    run:
        # 获取当前样本的数据类型
        current_datatype = sample_datatype_map[wildcards.name]
        # 把输入reads转成逗号分隔的字符串,适配你的脚本参数
        reads_str = ",".join(input)
        # 执行单样本的预处理命令
        shell(
            "python preprocess.py -i {reads} -t {threads} -o {out_dir} -adaptdir MOSCA/Databases/illumina_adapters -rrnadbs MOSCA/Databases/rRNA_databases -d {datatype}",
            reads=reads_str,
            threads=threads,
            out_dir=config["output"],
            datatype=current_datatype
        )

关键调整说明

  • 样本映射字典:我创建了sample_reads_map和sample_datatype_map两个字典,把每个样本名和它对应的输入文件、数据类型绑定起来。这样后续可以通过样本名快速查到单个样本的参数,不会再把所有样本的参数混在一起。
  • 通配符驱动任务:规则里用{name}作为通配符,Snakemake会自动遍历所有样本名,为每个样本创建独立的任务。wildcard_constraints还能防止出现无效的样本名任务。
  • 独立的输入输出:输入通过lambda函数根据当前任务的{name}获取对应样本的reads,输出也是针对单个样本生成的路径,每个样本的输出完全独立,不会互相干扰。
  • 单样本参数传递:在run块里,通过wildcards.name拿到当前任务的样本名,再从映射字典里取出对应的数据类型,确保每个任务只用到自己的参数。

额外小提示

  • 确保experiments['Name']中的每个样本名都是唯一的,如果你的命名逻辑可能产生重复,可以结合Sample和Condition列来生成更独特的名字,比如:
    experiments['Name'] = experiments['Sample'] + "_" + experiments['Condition'] + "_" + [filename.split('/')[-1].split('_R')[0] for filename in experiments['Files'].str.split(',').str[0]]
    
  • 如果存在单端测序的样本,可以新增一个sample_is_paired映射字典,在生成输出时动态调整fr的取值:
    sample_is_paired = dict(zip(experiments['Name'], experiments['Files'].str.contains(',')))
    # 然后在output里改成:
    fr=['_forward_paired', '_reverse_paired'] if sample_is_paired[wildcards.name] else ['']
    

内容的提问来源于stack exchange,提问作者João Sequeira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:42:42