Snakemake中如何正确指定glob_wildcards获取到的文件
问题解答
首先直接给出结论:你代码里直接写SAMPLES[0]、SAMPLES[1]的写法不正确,存在两个核心问题:
glob_wildcards返回的通配符匹配结果是无序的,没有任何规范保证返回列表的索引顺序,索引0不一定对应sample1.txt,索引1也不一定对应sample2.txt,换运行环境、文件系统时很容易出现匹配错误。SAMPLES列表里存储的只是从文件名中提取到的样本名字符串(比如sample1、sample2),不包含完整的文件路径和后缀,直接作为输入传入规则时,Snakemake无法定位到实际文件。
正确写法
首先先修正原代码的基础错误:你代码中expand("{sample}txt", sample=SAMPLES)缺少文件名的点分隔符,且和glob_wildcards匹配的data/目录路径不统一,需要先对齐路径规则。
要稳定指定拼接sample1.txt和sample2.txt,最稳妥的方式是明确指定目标样本名,通过expand生成完整文件路径,参考代码如下:
FILES = glob_wildcards("data/{sample}.txt") SAMPLES = FILES.sample # 明确声明需要拼接的目标样本,不依赖通配符返回顺序 CONCAT_TARGETS = ["sample1", "sample2"] rule all: input: expand("data/{sample}.txt", sample=SAMPLES), "data/concat.txt" rule concat: input: expand("data/{sample}.txt", sample=CONCAT_TARGETS) output: "data/concat.txt" shell: "cat {input} > {output}"
补充说明
- 如果你确实需要从
glob_wildcards自动获取的样本列表中筛选目标文件,不要靠索引取值,要按样本名做匹配筛选,避免顺序问题:# 按样本名匹配筛选,比索引取值可靠 sample1 = [s for s in SAMPLES if s == "sample1"][0] sample2 = [s for s in SAMPLES if s == "sample2"][0] - 所有传入规则的输入文件路径,要和
glob_wildcards匹配的路径规则保持一致,否则会出现文件找不到的报错。
内容的提问来源于stack exchange,提问作者mimi
相关产品推荐
相关产品推荐

