You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake规则传入文件路径列表并在Click脚本中读取的方法

Snakemake向Click脚本传入路径列表的实现方法

修正Snakemake规则代码

原代码存在三个问题需要调整:

  • 不能用占位字符串代替路径列表,需要先通过文件匹配拿到path_1下所有符合后缀要求的文件路径
  • 原shell块的两行命令没有分隔符,直接拼接会成为无效命令
  • Snakemake的input字段如果传入列表,在shell块中引用时会自动展开为空格分隔的字符串,可直接传给脚本

修正后的完整Snakemake代码:

import glob

path_1 = 'data/raw/data2process/'
path_2 = 'data/raw/table.xlsx'
# 收集目录下所有.xlsx/.csv/.xls后缀的文件,生成路径列表
list_of_pathes = []
for suffix in [".xlsx", ".csv", ".xls"]:
    list_of_pathes.extend(glob.glob(f"{path_1}*{suffix}"))

rule merge_tables:
    input:
        list_of_pathes = list_of_pathes,
        other_table = path_2
    output:
        res = "results/final_output.xlsx" # 替换为实际输出路径
    shell:
        """
        python script_1.py {input.list_of_pathes} {output.res}
        python script_2.py {input.other_table} {output.res}
        """

修正script_1.py的Click参数定义

原代码中Click参数没有指定接收多值的规则,这里用nargs=-1配置参数,接收任意数量的位置参数,再通过Python解包语法拆分输入路径列表和输出路径:

  • 所有传入的位置参数会被打包为元组
  • 元组最后一位为输出路径,前面所有元素为待处理的文件路径列表
  • 自动校验传入的文件路径是否存在,避免无效输入

修正后的script_1.py代码:

import click
import pandas as pd

@click.command()
# nargs=-1 表示接收任意个位置参数,required=True表示至少传1个参数
@click.argument("paths", nargs=-1, type=click.Path(exists=True), required=True)
def foo(paths):
    # 解包参数:最后一个值是输出路径,前面所有值是输入文件列表
    *input_list_of_pathes, out_path = paths
    df = pd.DataFrame()
    for path in input_list_of_pathes:
        # 按后缀选择读取方法,兼容csv和excel格式
        if path.endswith(".csv"):
            table = pd.read_csv(path)
        else:
            table = pd.read_excel(path)
        # 填入自定义处理逻辑
        df = pd.concat([df, table], ignore_index=True)
    df.to_excel(out_path, index=False)

if __name__ == "__main__":
    foo()

修正script_2.py的笔误

原代码中click.Path参数重复写了type=,且exist是拼写错误(正确为exists),修正后代码:

import click
import pandas as pd

@click.command()
@click.argument("input_path", type=click.Path(exists=True))
@click.argument("output_path", type=click.Path())
def foo_1(input_path: str, output_path: str):
    table = pd.read_excel(input_path)
    # 填入自定义处理逻辑
    table.to_excel(output_path, index=False)

if __name__ == "__main__":
    foo_1()

注意事项

  • 如果文件路径包含空格,Snakemake会自动给路径加转义引号,Click可正常识别,无需额外处理
  • 如果不想用位置参数解包的方式,也可将输入路径配置为multiple=True的Click选项,在Snakemake中拼接--input 路径1 --input 路径2格式传参即可,上述写法更简洁无需额外拼接
  • 运行时建议在shell命令中明确指定Python解释器,避免脚本无执行权限或解释器不匹配的问题

内容的提问来源于stack exchange,提问作者onetwoonexu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 06:19:28