You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake条件目标配置:实现样本与对应文件的匹配生成

解决Snakemake中样本与对应文件匹配的问题

问题背景

我有3个存储样本名的文件:

output/file_2.txt
output/file_4.txt
output/file_6.txt

每个文件内是对应所属的样本名,例如output/file_2.txt内容为:

sample1
sample2
sample3

原实现中,extract_ind()函数会读取所有文件的样本名合并为一个列表,在rule all里使用expand时,会生成n与sample的笛卡尔积,导致出现非匹配的组合(比如file_2对应来自file_4的sample4),不符合需求。

正确实现方法

方法1:预先生成目标文件路径列表

编写函数直接生成每个n对应的样本输出路径,避免笛卡尔积:

def get_target_files():
    targets = []
    for n in [2, 4, 6]:
        with open(f'output/file_{n}.txt') as f:
            samples = [line.strip() for line in f]
            targets.extend([f'output/file_{n}_{sample}.vcf.gz' for sample in samples])
    return targets

rule all:
    input:
        get_target_files()

方法2:利用expand的zip模式关联配对

优化你写的use_ind函数,结合expand的zip参数,让n和对应样本列表一一配对:

def use_ind(n):
    with open(f'output/file_{n}.txt') as f:
        return [line.strip() for line in f]

# 准备n列表和对应的样本列表
n_list = [2, 4, 6]
sample_groups = [use_ind(n) for n in n_list]

rule all:
    input:
        expand('output/file_{n}_{sample}.vcf.gz', zip, n=n_list, sample=sample_groups)

修正extract_samples规则的shell命令

原shell命令会遍历输入VCF的所有样本,应该直接使用wildcard中的sample参数,避免多余操作:

rule extract_samples:
    input:
        vcf = 'muscle/file_{n}.vcf.gz'
    output:
        vcf_dir = 'output/file_{n}_{sample}.vcf.gz'
    shell:
        'bcftools view -O z -s {wildcards.sample} -o {output.vcf_dir} {input.vcf}'

内容的提问来源于stack exchange,提问作者moth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 13:24:21