You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake使用directory输出文件时通配符错误排查

问题描述

作为Snakemake新手,我尝试在规则中使用另一个克隆Git仓库规则的directory输出文件,但报错:Wildcards in input files cannot be determined from output files: 'json_file'。

我已经完成Carpentries的Snakemake教程,我的工作流和教程的区别在于:教程中数据已存在,而我需要第一步生成后续使用的数据。工作流逻辑如下:

  • 将Git仓库克隆到路径{path};
  • 并行运行脚本{script}处理{path}/parsed/下的所有JSON文件,生成聚合结果{result}。

目前只有clone_git规则能正常运行(把rule all的input设为GIT_PATH时),我怀疑是JSON文件在工作流启动时还不存在导致的错误,而且这个流程是通过module引入的子工作流。

附相关代码:

GIT_PATH = config['git_local_path']  # git/
PARSED_JSON_PATH = f'{GIT_PATH}parsed/'
GIT_URL = config['git_url']

# A single parsed JSON file
PARSED_JSON_FILE = f'{PARSED_JSON_PATH}{{json_file}}.json'

# Build a list of parsed JSON file names
PARSED_JSON_FILE_NAMES = glob_wildcards(PARSED_JSON_FILE).json_file

# All parsed JSON files
ALL_PARSED_JSONS = expand(PARSED_JSON_FILE, json_file=PARSED_JSON_FILE_NAMES)


rule all:
    input: 'result.json'

rule clone_git:
    output: directory(GIT_PATH)
    threads: 1
    conda: f'{ENVS_DIR}git.yml'
    shell: f'git clone --depth 1 {GIT_URL} {{output}}'

rule extract_json:
    input:
        cmd='scripts/extract_json.py',
        json_file=PARSED_JSON_FILE
    output: 'result.json'
    threads: 50
    shell: 'python {input.cmd} {input.json_file} {output}'
问题原因与解决方案

核心原因

  1. 提前执行的文件匹配无效:你用glob_wildcards在工作流启动阶段就去匹配PARSED_JSON_PATH下的文件,但此时Git仓库还没克隆,目标目录是空的,导致PARSED_JSON_FILE_NAMES为空,后续通配符解析失效。
  2. 通配符与输出不匹配:extract_json规则的输入带json_file通配符,但输出是单个文件result.json,Snakemake无法从单个输出反向推断出多个输入的通配符值,直接触发报错。

解决方法(两种可选)

方法一:直接批量处理(适合脚本支持批量读取文件)

如果你的extract_json.py本身可以接受目录或多个文件路径作为输入,直接让规则依赖克隆完成的仓库,然后在shell命令中批量传入JSON文件:

GIT_PATH = config['git_local_path']  # git/
PARSED_JSON_PATH = f'{GIT_PATH}parsed/'
GIT_URL = config['git_url']

rule all:
    input: 'result.json'

rule clone_git:
    output: directory(GIT_PATH)
    threads: 1
    conda: f'{ENVS_DIR}git.yml'
    shell: f'git clone --depth 1 {GIT_URL} {{output}}'

rule extract_json:
    input:
        cmd='scripts/extract_json.py',
        # 依赖克隆完成的仓库,确保JSON文件已存在
        git_repo=directory(GIT_PATH)
    output: 'result.json'
    threads: 50
    # 直接传入目录下所有JSON文件给脚本
    shell: 'python {input.cmd} {PARSED_JSON_PATH}*.json {output}'

方法二:拆分规则并行处理单个文件(适合需要并行优化)

如果JSON文件数量多,需要并行处理,先拆分出单个文件的处理规则,再汇总结果:

GIT_PATH = config['git_local_path']  # git/
PARSED_JSON_PATH = f'{GIT_PATH}parsed/'
GIT_URL = config['git_url']

rule all:
    input: 'result.json'

rule clone_git:
    output: directory(GIT_PATH)
    threads: 1
    conda: f'{ENVS_DIR}git.yml'
    shell: f'git clone --depth 1 {GIT_URL} {{output}}'

# 并行处理单个JSON文件,生成临时结果
rule process_single_json:
    input:
        cmd='scripts/process_single.py',  # 处理单个JSON的脚本
        json_file=os.path.join(PARSED_JSON_PATH, '{json_file}.json'),
        git_repo=directory(GIT_PATH)
    output: 'temp_{json_file}.json'
    shell: 'python {input.cmd} {input.json_file} {output}'

# 汇总所有临时结果生成最终文件
rule extract_json:
    input:
        cmd='scripts/extract_json.py',
        # 克隆完成后再获取JSON文件列表
        temp_results=expand(
            'temp_{json_file}.json',
            json_file=glob_wildcards(os.path.join(PARSED_JSON_PATH, '{json_file}.json')).json_file
        )
    output: 'result.json'
    threads: 50
    shell: 'python {input.cmd} {input.temp_results} {output}'

关键说明

  • 无论哪种方法,核心都是确保Git克隆规则先执行,后续规则才能获取到生成的JSON文件。
  • 避免在工作流初始化阶段(规则外)去匹配动态生成的文件,这类操作要放到规则内部或依赖完成后执行。

内容的提问来源于stack exchange,提问作者QueNuevo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 17:40:21