Snakemake分规则解压缩与处理文件报错,求可行解决方案
Snakemake分离规则执行MissingInputException问题解决
问题场景
我是Snakemake初级用户,想用两个独立规则分别实现文件解压缩和数据处理,但Part1的分离规则执行时抛出MissingInputException错误。虽通过合并规则的Part2实现了功能,但希望用分离规则的方式解决,想知道Snakemake有没有内置方法让Part1生效,同时需了解问题原因和解决原理。
复现代码
# 导入必要模块 import os import numpy as np import pandas as pd import zipfile import shutil # 创建所需目录和文件 os.makedirs("path/to/my/file/", exist_ok=True) os.makedirs("path/to/my/compressed", exist_ok=True) os.makedirs("path/to/my/uncompressed", exist_ok=True) # 生成随机数据并创建DataFrame x = np.random.rand(3, 2) df = pd.DataFrame(data=x.astype(float)) df.to_csv("path/to/my/file/data.csv") # 将CSV文件压缩为zip包 with zipfile.ZipFile('path/to/my/compressed/myfile.zip', 'w') as z: z.write("path/to/my/file/data.csv", "data.csv") # 主规则(rule all)指定预期输出文件 rule all: input: i1='path/to/my/uncompressed', i2='path/to/my/Routput/data.csv' # Part 1: 报错的分离规则 # 第一个规则解压compressed目录下的文件到uncompressed目录 # 第二个规则读取解压后的文件用R处理后写入Routput目录 # 此写法报错:MissingInputException at line 37 of the Snakefile rule uncompress: input: 'path/to/my/compressed/myfile.zip' output: directory('path/to/my/uncompressed') shell: """ unzip {input} -d {output} """ rule load_data: input: 'path/to/my/uncompressed/data.csv' output: 'path/to/my/Routput/data.csv' shell: """ Rscript -e "x={input}; y={output}; X=read.csv(x, header=T, sep=','); write.csv(X, y)" """ # Part 2: 可行但不推荐的合并规则方案 # 注释Part1并取消注释Part2可执行 # 将数据处理逻辑嵌入解压规则,虽能运行但不符合分离规则的需求 # rule uncompress: # input: # 'path/to/my/compressed/myfile.zip' # output: # o1=directory('path/to/my/uncompressed'), # o2='path/to/my/Routput/data.csv' # params: # p1='path/to/my/uncompressed/data.csv' # shell: # """ # unzip {input} -d {output.o1} # Rscript -e "x='{params.p1}' ;y= '{output.o2}'; X=read.csv(x, header=T, sep=','); write.csv(X, file=y)" # """
环境信息
- Snakemake版本:5.10.0
- Python版本:3.8.10
- Linux发行版:Ubuntu 20.04.6 LTS
问题原因
核心问题是**directory()输出无法关联目录内的具体文件**:
uncompress规则仅声明directory('path/to/my/uncompressed')为输出,Snakemake只会校验该目录是否存在,不会追踪目录内生成的data.csv。- 当
load_data规则依赖path/to/my/uncompressed/data.csv时,Snakemake找不到任何规则明确声明会生成这个文件——即便uncompress实际会生成它,但规则输出未明确标注,Snakemake的依赖解析系统无法建立两者的关联,因此抛出输入缺失的错误。
分离规则生效的解决方法
方法1:明确声明解压后的具体文件作为输出
修改uncompress规则,直接把解压得到的具体文件列为输出,而非仅声明目录:
rule uncompress: input: 'path/to/my/compressed/myfile.zip' output: 'path/to/my/uncompressed/data.csv' # 明确指定输出的具体文件 shell: """ unzip {input} -d $(dirname {output}) # 用dirname获取文件所在目录 """ rule load_data: input: 'path/to/my/uncompressed/data.csv' output: 'path/to/my/Routput/data.csv' shell: """ Rscript -e "x='{input}'; y='{output}'; X=read.csv(x, header=T, sep=','); write.csv(X, y)" """
方法2:通配符+expand(适配批量解压场景)
如果需要处理多个文件,可通过通配符建立规则间的关联:
rule all: input: expand('path/to/my/uncompressed/{file}.csv', file=['data']), 'path/to/my/Routput/data.csv' rule uncompress: input: 'path/to/my/compressed/myfile.zip' output: 'path/to/my/uncompressed/{file}.csv' shell: """ unzip {input} {wildcards.file}.csv -d $(dirname {output}) """ rule load_data: input: 'path/to/my/uncompressed/{file}.csv' output: 'path/to/my/Routput/{file}.csv' shell: """ Rscript -e "x='{input}'; y='{output}'; X=read.csv(x, header=T, sep=','); write.csv(X, y)" """
解决原理
Snakemake的依赖解析基于具体文件的显式声明,而非对目录内容的隐式感知:
- 规则的
output必须明确列出所有需要被后续规则依赖的文件,这样Snakemake才能构建正确的依赖关系图(DAG)。 - 仅声明
directory()时,Snakemake仅验证目录存在性,不追踪目录内的文件生成或变化,后续规则依赖目录内文件时,Snakemake无法识别该文件由前置规则生成,因此判定输入缺失。
内容的提问来源于stack exchange,提问作者Sam
相关产品推荐
相关产品推荐

