You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Snakemake中实现参数的条件组合?

解决Snakemake中样本与lane的依赖匹配问题

我有一批带lane(L0*)、read(Read*)、sample(XY*)标识的FASTQ文件,所有样本都包含Read1和Read2,多数样本仅属于一个lane,但部分样本跨多个lane。使用Snakemake v5.26.1构建工作流时,直接用expand生成输入输出路径会产生无效的sample-lane组合,触发MissingInputException报错。需要实现lane值依赖于sample值的精准匹配,也可考虑升级版本。


输入文件列表

$ ls -1 inputs/ 
CM0009619-ABB_L01_Read1_Sample_Library_XY102.fastq.gz
CM0009619-ABB_L01_Read1_Sample_Library_XY110.fastq.gz
CM0009619-ABB_L01_Read1_Sample_Library_XY84.fastq.gz
CM0009619-ABB_L01_Read2_Sample_Library_XY102.fastq.gz
CM0009619-ABB_L01_Read2_Sample_Library_XY110.fastq.gz
CM0009619-ABB_L01_Read2_Sample_Library_XY84.fastq.gz
CM0009619-ABB_L02_Read1_Sample_Library_XY84.fastq.gz
CM0009619-ABB_L02_Read1_Sample_Library_XY88.fastq.gz
CM0009619-ABB_L02_Read2_Sample_Library_XY84.fastq.gz
CM0009619-ABB_L02_Read2_Sample_Library_XY88.fastq.gz

错误的Snakefile代码

SAMPLES = ["XY84", "XY88", "XY102", "XY110"]

rule copy_files:
  input:
    expand("inputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz", sample=SAMPLES, read=[1,2], lane=[1,2] )   
  output:
    expand("outputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz", sample=SAMPLES, read=[1,2], lane=[1,2] )
  shell:
    "cp {input} {output}"   

报错信息

$ snakemake --snakefile Snakefile --cores 1
Building DAG of jobs...
MissingInputException in line 10 of /home/UNIXHOME/sholt/smktests/Snakefile:
Missing input files for rule copy_files:
inputs/CM0009619-ABB_L02_Read2_Sample_Library_XY102.fastq.gz
inputs/CM0009619-ABB_L01_Read1_Sample_Library_XY88.fastq.gz
inputs/CM0009619-ABB_L02_Read1_Sample_Library_XY102.fastq.gz
inputs/CM0009619-ABB_L02_Read1_Sample_Library_XY110.fastq.gz
inputs/CM0009619-ABB_L01_Read2_Sample_Library_XY88.fastq.gz
inputs/CM0009619-ABB_L02_Read2_Sample_Library_XY110.fastq.gz

解决方案

方法1:手动定义样本与lane的映射字典

直接明确每个样本对应的lane列表,通过zip参数绑定组合关系,避免无效匹配:

# 定义样本到lane的映射关系
SAMPLE_LANES = {
    "XY84": ["1", "2"],
    "XY88": ["2"],
    "XY102": ["1"],
    "XY110": ["1"]
}

# 生成有效的(sample, lane)组合对
SAMPLE_LANE_PAIRS = [(s, l) for s, lanes in SAMPLE_LANES.items() for l in lanes]

rule copy_files:
  input:
    expand("inputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz", 
           read=[1,2],
           zip,  # 绑定sample和lane的对应关系
           sample=[pair[0] for pair in SAMPLE_LANE_PAIRS],
           lane=[pair[1] for pair in SAMPLE_LANE_PAIRS]
           )
  output:
    expand("outputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz", 
           read=[1,2],
           zip,
           sample=[pair[0] for pair in SAMPLE_LANE_PAIRS],
           lane=[pair[1] for pair in SAMPLE_LANE_PAIRS]
           )
  shell:
    "cp {input} {output}"

方法2:从输入文件自动提取有效组合(更灵活)

若样本和lane组合较多,手动维护字典麻烦,可通过扫描输入文件自动提取有效组合:

import glob
import re

# 扫描输入文件,提取所有有效的sample和lane组合
input_files = glob.glob("inputs/*.fastq.gz")
sample_lane_map = {}
pattern = re.compile(r"_L0(\d+)_.*_XY(\d+)\.fastq\.gz")

for f in input_files:
    match = pattern.search(f)
    if match:
        lane = match.group(1)
        sample = f"XY{match.group(2)}"
        if sample not in sample_lane_map:
            sample_lane_map[sample] = set()
        sample_lane_map[sample].add(lane)

# 转换为有序组合对
SAMPLE_LANE_PAIRS = [(s, l) for s, lanes in sample_lane_map.items() for l in sorted(lanes)]

rule copy_files:
  input:
    expand("inputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz", 
           read=[1,2],
           zip,
           sample=[pair[0] for pair in SAMPLE_LANE_PAIRS],
           lane=[pair[1] for pair in SAMPLE_LANE_PAIRS]
           )
  output:
    expand("outputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz", 
           read=[1,2],
           zip,
           sample=[pair[0] for pair in SAMPLE_LANE_PAIRS],
           lane=[pair[1] for pair in SAMPLE_LANE_PAIRS]
           )
  shell:
    "cp {input} {output}"

方法3:拆分规则为单样本单lane任务(最佳实践)

将大规则拆分为每个样本+lane+read的独立任务,让Snakemake自动处理依赖,逻辑更清晰且易于扩展:

import glob
import re

input_files = glob.glob("inputs/*.fastq.gz")
pattern = re.compile(r"_L0(\d+)_Read(\d+)_.*_XY(\d+)\.fastq\.gz")

# 提取所有有效任务参数
TASKS = []
for f in input_files:
    match = pattern.search(f)
    if match:
        lane = match.group(1)
        read = match.group(2)
        sample = f"XY{match.group(3)}"
        TASKS.append( (sample, lane, read) )

rule all:
    input:
        expand("outputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz",
               zip,
               sample=[t[0] for t in TASKS],
               lane=[t[1] for t in TASKS],
               read=[t[2] for t in TASKS]
               )

rule copy_file:
    input:
        "inputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz"
    output:
        "outputs/CM0009619-ABB_L0{lane}_Read{read}_Sample_Library_{sample}.fastq.gz"
    shell:
        "cp {input} {output}"

版本兼容性说明

  • 上述方法在v5.26.1版本均可正常工作,zip参数是Snakemake原生支持的功能,无需升级。
  • 若后续需要更复杂的参数展开逻辑,升级到最新稳定版本(如v7+)可获得更多高级功能(如expand的wildcards参数、更灵活的规则模板),但当前问题用现有版本完全可以解决。

内容的提问来源于stack exchange,提问作者Sam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 21:55:12