snakedeploy远程部署的GitHub工作流访问自有目录文件的最佳实践
Snakedeploy远程工作流访问自身资源的最佳实践
方案合法性确认
你梳理的infer_source_file()判断文件类型 + workflow.sourcecache.open()处理文件的方案就是Snakemake官方指定的标准实现。该能力从6.8.1版本推出,就是为了解决远程模块场景下srcdir()返回URL无法直接读取的问题,早期确实存在官方文档更新滞后的情况,目前最新稳定版文档已经补充了相关API的说明。
这套方案自带缓存机制,同一版本的远程资源只会拉取一次,并且和Snakemake内部的哈希校验逻辑打通,能避免资源版本不一致的问题。
四类需求的具体实现示例
你提到的四类使用场景可以参考以下代码实现:
- 运行
workflow/scripts/*.py中的脚本
对于需要作为命令执行的脚本,先通过sourcecache获取本地缓存路径再调用:from snakemake.utils import infer_source_file # 解析脚本的源路径,自动适配本地和远程场景 script_path = infer_source_file("workflow/scripts/analysis.py") # 拉取远程文件到本地缓存(本地文件直接返回路径) local_script = workflow.sourcecache.get(script_path) rule run_analysis: input: "data/input.txt" output: "results/output.txt" shell: "python {local_script} {input} {output}" - 加载
resources/<virus>/<reference>.fasta序列、workflow/schemas/*.jsonschema文件、额外YAML配置
这三类需要读取内容的文件可直接用workflow.sourcecache.open()打开处理:import json import yaml from snakemake.utils import infer_source_file # 读取参考基因组 ref_path = infer_source_file(f"resources/{virus}/{reference}.fasta") with workflow.sourcecache.open(ref_path) as f: ref_seq = f.read() # 读取JSON schema schema_path = infer_source_file("workflow/schemas/config.schema.json") with workflow.sourcecache.open(schema_path) as f: schema = json.load(f) # 读取病毒专属配置 hiv_config_path = infer_source_file("config/hiv.yaml") with workflow.sourcecache.open(hiv_config_path) as f: hiv_config = yaml.safe_load(f)
其他可选最佳实践
如果你的工作流仅适配Snakemake 7.12及以上版本,可以使用更简化的写法,无需手动调用infer_source_file:
# 7.12+ 版本简化写法,直接获取当前模块的源目录 module_srcdir = workflow.modules.current.srcdir local_script = workflow.sourcecache.get(module_srcdir / "scripts/analysis.py")
如果工作流包含大量大体积静态资源(比如参考基因组),建议将这类资源放到独立的资源仓库,通过Snakemake的remote()接口单独托管,避免工作流仓库体积过大,降低部署时的拉取速度。
内容的提问来源于stack exchange,提问作者DrYak
相关产品推荐
相关产品推荐

