Snakemake规则:已知数量可变输出文件的实现难题
解决Snakemake可变输出归档解压的方案
针对你这种已知归档内文件数量、无法在output中使用函数,且规则运行成本极高不能多次执行的场景,我给你几个实用的替代方案:
方案1:通过配置文件映射归档与输出文件
这是最直接的方式,利用运行前已知的文件数量,在配置文件里提前定义每个归档对应的输出文件列表,然后在规则里用expand来生成明确的输出。
首先创建一个config.yaml文件,把每个归档和它的输出文件对应起来:
archives: data_sample.tar.gz: outputs: - unpacked/sample/file1.txt - unpacked/sample/file2.csv - unpacked/sample/file3.log raw_data.zip: outputs: - unpacked/raw/data.img - unpacked/raw/metadata.json
然后在你的Snakefile里这样写规则:
configfile: "config.yaml" rule unpack_archive: input: lambda wildcards: wildcards.archive # 匹配归档文件名 output: expand("{out}", out=config["archives"][wildcards.archive]["outputs"]) shell: """ # 根据归档类型选择解压命令,确保文件输出到指定路径 ARCHIVE_DIR=$(dirname {output[0]}) mkdir -p $ARCHIVE_DIR if [[ "{input}" == *.tar.gz ]]; then tar -xzf {input} -C $ARCHIVE_DIR elif [[ "{input}" == *.zip ]]; then unzip {input} -d $ARCHIVE_DIR fi """
这个方案完全避免了在output里使用自定义函数,而是用提前配置好的静态文件列表,Snakemake能明确识别输出,且每个归档只运行一次规则。
方案2:在Snakefile中直接定义归档-输出映射
如果不想用外部配置文件,也可以直接在Snakefile开头定义一个字典,把归档和对应的输出文件列表关联起来:
# 提前定义每个归档对应的输出文件 ARCHIVE_OUTPUT_MAP = { "data_sample.tar.gz": [ "unpacked/sample/file1.txt", "unpacked/sample/file2.csv", "unpacked/sample/file3.log" ], "raw_data.zip": [ "unpacked/raw/data.img", "unpacked/raw/metadata.json" ] } rule unpack_archive: input: lambda wc: wc.archive output: expand("{file}", file=ARCHIVE_OUTPUT_MAP[wc.archive]) shell: """ ARCHIVE_DIR=$(dirname {output[0]}) mkdir -p $ARCHIVE_DIR # 解压命令同方案1 if [[ "{input}" == *.tar.gz ]]; then tar -xzf {input} -C $ARCHIVE_DIR elif [[ "{input}" == *.zip ]]; then unzip {input} -d $ARCHIVE_DIR fi """
这个方案和方案1逻辑一致,只是把映射关系放在了Snakefile内部,适合小型项目或不想维护额外配置文件的情况。
方案3:使用Checkpoint处理动态输出(灵活备选)
如果虽然已知文件数量,但不想硬编码所有输出路径,可以用Snakemake的Checkpoint功能。它允许规则输出一个目录,后续再收集目录内的文件,同样能保证每个归档只运行一次解压规则:
checkpoint unpack_archive: input: "{archive}" output: directory("unpacked/{archive}_dir") # 输出一个目录 shell: """ mkdir -p {output} # 解压到指定目录 if [[ "{input}" == *.tar.gz ]]; then tar -xzf {input} -C {output} elif [[ "{input}" == *.zip ]]; then unzip {input} -d {output} fi """ # 后续需要使用这些文件的规则,可以通过checkpoint_output收集 rule process_unpacked_files: input: # 收集checkpoint输出目录下的所有文件 lambda wildcards: checkpoint_output("unpack_archive", archive=wildcards.archive).glob("*") output: "processed/{archive}_done.txt" shell: """ # 这里写你的文件处理逻辑 echo "Processed all files from {wildcards.archive}" > {output} """
这个方案不需要提前写死所有输出文件路径,但需要后续规则来收集输出,适合归档内文件路径有规律但不想逐个列出来的场景。
这些方案都能满足你“一次运行规则、不使用output函数、利用已知文件数量配置”的需求,你可以根据自己的项目规模和偏好选择~
内容的提问来源于stack exchange,提问作者Dion
相关产品推荐
相关产品推荐

