You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake:无需指定列表聚合time_average文件的最佳实践

问题描述

我的目录结构如下:

-- path
   -- parameter_combination_1
      - time_average.property1.csv
      - time_average.property2.csv
      - ...
   -- parameter_combination_2
      - time_average.property1.csv
      - time_average.property2.csv
      - ...
   -- ...

我想创建一条Snakemake规则,针对通配符{filename}(例如property1.csv)和{path},聚合所有名称包含time_average的文件。对应输入文件示例:

  • path/parameter_combination_1/time_average.property1.csv
  • path/parameter_combination_2/time_average.property1.csv
  • path/parameter_combination_3/time_average.property1.csv
  • ...

我知道可以用expand指定固定参数组合列表,比如:

rule aggregate_time_averages:
    input:
        expand("{path}/{parameter_combination}/time_average.{filename}", parameter_combination=['parameter_combination_1', 'parameter_combination_2'])
    output:
        "{path}/aggregate.{filename}"

但我不想手动指定固定列表,希望能自动匹配所有parameter_combination_*文件夹。

我尝试过全局使用glob_wildcards:

PROJECT_PATHS, PARAMS_COMBS, SEEDS = glob_wildcards("{path}/{parameter_combination}/{seed}/config.yaml")

rule aggregate_time_averages:
    input:
        expand("{path}/{parameter_combination}/time_avg.{filename}", parameter_combination=PARAMS_COMBS)
    output:
        "{path}/aggregate.{filename}"

执行命令snakemake --cores 1 test_project/aggregate.order_parameter.csv --use-conda时,报错No values given for wildcard 'path'。而且全局glob_wildcards会获取所有匹配的通配符,我需要的是仅匹配当前规则调用的{path}/{filename}对应的parameter_combination,希望能在规则内完成匹配。

解决方案

针对这个场景,推荐以下两种最佳实践:

1. 规则内动态调用glob_wildcards

利用Snakemake的输入函数特性,在规则内部针对当前的{path}和{filename}动态匹配对应参数组合文件夹,避免全局扫描的冗余和冲突:

rule aggregate_time_averages:
    output:
        "{path}/aggregate.{filename}"
    input:
        lambda wildcards: expand(
            "{path}/{param_comb}/time_average.{filename}",
            path=wildcards.path,
            filename=wildcards.filename,
            param_comb=glob_wildcards(
                "{path}/{param_comb}/time_average.{filename}",
                path=wildcards.path,
                filename=wildcards.filename
            ).param_comb
        )
    shell:
        # 替换为你的聚合逻辑,例如合并CSV文件
        "cat {input} > {output}"

核心逻辑:

  • 通过lambda wildcards传入当前规则的通配符实际值
  • 限定glob_wildcards仅扫描当前path和filename对应的参数组合文件夹
  • 用expand生成完整的输入文件路径列表

2. 使用dynamic输入(Snakemake 7.0+)

如果你的Snakemake版本在7.0及以上,可以用dynamic语法自动匹配文件,写法更简洁:

rule aggregate_time_averages:
    output:
        "{path}/aggregate.{filename}"
    input:
        dynamic("{path}/parameter_combination_*/time_average.{filename}")
    shell:
        "cat {input} > {output}"

dynamic会在规则执行前自动扫描匹配的文件,无需手动处理通配符。注意需确保参数组合文件夹严格遵循parameter_combination_*的命名格式。

关键注意事项

  • 检查路径模板与实际文件的一致性:比如之前代码中time_avg和time_average的拼写差异会导致匹配失败,需完全对应
  • 若使用dynamic,需确认Snakemake版本满足要求,旧版本不支持该特性

内容的提问来源于stack exchange,提问作者zanzu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 10:45:27