Snakemake:无需指定列表聚合time_average文件的最佳实践
问题描述
我的目录结构如下:
-- path -- parameter_combination_1 - time_average.property1.csv - time_average.property2.csv - ... -- parameter_combination_2 - time_average.property1.csv - time_average.property2.csv - ... -- ...
我想创建一条Snakemake规则,针对通配符{filename}(例如property1.csv)和{path},聚合所有名称包含time_average的文件。对应输入文件示例:
path/parameter_combination_1/time_average.property1.csvpath/parameter_combination_2/time_average.property1.csvpath/parameter_combination_3/time_average.property1.csv- ...
我知道可以用expand指定固定参数组合列表,比如:
rule aggregate_time_averages: input: expand("{path}/{parameter_combination}/time_average.{filename}", parameter_combination=['parameter_combination_1', 'parameter_combination_2']) output: "{path}/aggregate.{filename}"
但我不想手动指定固定列表,希望能自动匹配所有parameter_combination_*文件夹。
我尝试过全局使用glob_wildcards:
PROJECT_PATHS, PARAMS_COMBS, SEEDS = glob_wildcards("{path}/{parameter_combination}/{seed}/config.yaml") rule aggregate_time_averages: input: expand("{path}/{parameter_combination}/time_avg.{filename}", parameter_combination=PARAMS_COMBS) output: "{path}/aggregate.{filename}"
执行命令snakemake --cores 1 test_project/aggregate.order_parameter.csv --use-conda时,报错No values given for wildcard 'path'。而且全局glob_wildcards会获取所有匹配的通配符,我需要的是仅匹配当前规则调用的{path}/{filename}对应的parameter_combination,希望能在规则内完成匹配。
解决方案
针对这个场景,推荐以下两种最佳实践:
1. 规则内动态调用glob_wildcards
利用Snakemake的输入函数特性,在规则内部针对当前的{path}和{filename}动态匹配对应参数组合文件夹,避免全局扫描的冗余和冲突:
rule aggregate_time_averages: output: "{path}/aggregate.{filename}" input: lambda wildcards: expand( "{path}/{param_comb}/time_average.{filename}", path=wildcards.path, filename=wildcards.filename, param_comb=glob_wildcards( "{path}/{param_comb}/time_average.{filename}", path=wildcards.path, filename=wildcards.filename ).param_comb ) shell: # 替换为你的聚合逻辑,例如合并CSV文件 "cat {input} > {output}"
核心逻辑:
- 通过
lambda wildcards传入当前规则的通配符实际值 - 限定
glob_wildcards仅扫描当前path和filename对应的参数组合文件夹 - 用
expand生成完整的输入文件路径列表
2. 使用dynamic输入(Snakemake 7.0+)
如果你的Snakemake版本在7.0及以上,可以用dynamic语法自动匹配文件,写法更简洁:
rule aggregate_time_averages: output: "{path}/aggregate.{filename}" input: dynamic("{path}/parameter_combination_*/time_average.{filename}") shell: "cat {input} > {output}"
dynamic会在规则执行前自动扫描匹配的文件,无需手动处理通配符。注意需确保参数组合文件夹严格遵循parameter_combination_*的命名格式。
关键注意事项
- 检查路径模板与实际文件的一致性:比如之前代码中
time_avg和time_average的拼写差异会导致匹配失败,需完全对应 - 若使用
dynamic,需确认Snakemake版本满足要求,旧版本不支持该特性
内容的提问来源于stack exchange,提问作者zanzu
相关产品推荐
相关产品推荐

