You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake中如何正确指定glob_wildcards获取到的文件

问题解答

首先直接给出结论:你代码里直接写SAMPLES[0]、SAMPLES[1]的写法不正确,存在两个核心问题:

  1. glob_wildcards返回的通配符匹配结果是无序的,没有任何规范保证返回列表的索引顺序,索引0不一定对应sample1.txt,索引1也不一定对应sample2.txt,换运行环境、文件系统时很容易出现匹配错误。
  2. SAMPLES列表里存储的只是从文件名中提取到的样本名字符串(比如sample1、sample2),不包含完整的文件路径和后缀,直接作为输入传入规则时,Snakemake无法定位到实际文件。

正确写法

首先先修正原代码的基础错误:你代码中expand("{sample}txt", sample=SAMPLES)缺少文件名的点分隔符,且和glob_wildcards匹配的data/目录路径不统一,需要先对齐路径规则。
要稳定指定拼接sample1.txt和sample2.txt,最稳妥的方式是明确指定目标样本名,通过expand生成完整文件路径,参考代码如下:

FILES = glob_wildcards("data/{sample}.txt")
SAMPLES = FILES.sample

# 明确声明需要拼接的目标样本,不依赖通配符返回顺序
CONCAT_TARGETS = ["sample1", "sample2"]

rule all:
    input:
        expand("data/{sample}.txt", sample=SAMPLES),
        "data/concat.txt"

rule concat:
    input:
        expand("data/{sample}.txt", sample=CONCAT_TARGETS)
    output:
        "data/concat.txt"
    shell:
        "cat {input} > {output}"

补充说明

  • 如果你确实需要从glob_wildcards自动获取的样本列表中筛选目标文件,不要靠索引取值,要按样本名做匹配筛选,避免顺序问题:
    # 按样本名匹配筛选,比索引取值可靠
    sample1 = [s for s in SAMPLES if s == "sample1"][0]
    sample2 = [s for s in SAMPLES if s == "sample2"][0]
    
  • 所有传入规则的输入文件路径,要和glob_wildcards匹配的路径规则保持一致,否则会出现文件找不到的报错。

内容的提问来源于stack exchange,提问作者mimi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.01 02:57:21