Snakemake使用directory输出文件时通配符错误排查
问题描述
作为Snakemake新手,我尝试在规则中使用另一个克隆Git仓库规则的directory输出文件,但报错:Wildcards in input files cannot be determined from output files: 'json_file'。
我已经完成Carpentries的Snakemake教程,我的工作流和教程的区别在于:教程中数据已存在,而我需要第一步生成后续使用的数据。工作流逻辑如下:
- 将Git仓库克隆到路径
{path}; - 并行运行脚本
{script}处理{path}/parsed/下的所有JSON文件,生成聚合结果{result}。
目前只有clone_git规则能正常运行(把rule all的input设为GIT_PATH时),我怀疑是JSON文件在工作流启动时还不存在导致的错误,而且这个流程是通过module引入的子工作流。
附相关代码:
GIT_PATH = config['git_local_path'] # git/ PARSED_JSON_PATH = f'{GIT_PATH}parsed/' GIT_URL = config['git_url'] # A single parsed JSON file PARSED_JSON_FILE = f'{PARSED_JSON_PATH}{{json_file}}.json' # Build a list of parsed JSON file names PARSED_JSON_FILE_NAMES = glob_wildcards(PARSED_JSON_FILE).json_file # All parsed JSON files ALL_PARSED_JSONS = expand(PARSED_JSON_FILE, json_file=PARSED_JSON_FILE_NAMES) rule all: input: 'result.json' rule clone_git: output: directory(GIT_PATH) threads: 1 conda: f'{ENVS_DIR}git.yml' shell: f'git clone --depth 1 {GIT_URL} {{output}}' rule extract_json: input: cmd='scripts/extract_json.py', json_file=PARSED_JSON_FILE output: 'result.json' threads: 50 shell: 'python {input.cmd} {input.json_file} {output}'
问题原因与解决方案
核心原因
- 提前执行的文件匹配无效:你用
glob_wildcards在工作流启动阶段就去匹配PARSED_JSON_PATH下的文件,但此时Git仓库还没克隆,目标目录是空的,导致PARSED_JSON_FILE_NAMES为空,后续通配符解析失效。 - 通配符与输出不匹配:
extract_json规则的输入带json_file通配符,但输出是单个文件result.json,Snakemake无法从单个输出反向推断出多个输入的通配符值,直接触发报错。
解决方法(两种可选)
方法一:直接批量处理(适合脚本支持批量读取文件)
如果你的extract_json.py本身可以接受目录或多个文件路径作为输入,直接让规则依赖克隆完成的仓库,然后在shell命令中批量传入JSON文件:
GIT_PATH = config['git_local_path'] # git/ PARSED_JSON_PATH = f'{GIT_PATH}parsed/' GIT_URL = config['git_url'] rule all: input: 'result.json' rule clone_git: output: directory(GIT_PATH) threads: 1 conda: f'{ENVS_DIR}git.yml' shell: f'git clone --depth 1 {GIT_URL} {{output}}' rule extract_json: input: cmd='scripts/extract_json.py', # 依赖克隆完成的仓库,确保JSON文件已存在 git_repo=directory(GIT_PATH) output: 'result.json' threads: 50 # 直接传入目录下所有JSON文件给脚本 shell: 'python {input.cmd} {PARSED_JSON_PATH}*.json {output}'
方法二:拆分规则并行处理单个文件(适合需要并行优化)
如果JSON文件数量多,需要并行处理,先拆分出单个文件的处理规则,再汇总结果:
GIT_PATH = config['git_local_path'] # git/ PARSED_JSON_PATH = f'{GIT_PATH}parsed/' GIT_URL = config['git_url'] rule all: input: 'result.json' rule clone_git: output: directory(GIT_PATH) threads: 1 conda: f'{ENVS_DIR}git.yml' shell: f'git clone --depth 1 {GIT_URL} {{output}}' # 并行处理单个JSON文件,生成临时结果 rule process_single_json: input: cmd='scripts/process_single.py', # 处理单个JSON的脚本 json_file=os.path.join(PARSED_JSON_PATH, '{json_file}.json'), git_repo=directory(GIT_PATH) output: 'temp_{json_file}.json' shell: 'python {input.cmd} {input.json_file} {output}' # 汇总所有临时结果生成最终文件 rule extract_json: input: cmd='scripts/extract_json.py', # 克隆完成后再获取JSON文件列表 temp_results=expand( 'temp_{json_file}.json', json_file=glob_wildcards(os.path.join(PARSED_JSON_PATH, '{json_file}.json')).json_file ) output: 'result.json' threads: 50 shell: 'python {input.cmd} {input.temp_results} {output}'
关键说明
- 无论哪种方法,核心都是确保Git克隆规则先执行,后续规则才能获取到生成的JSON文件。
- 避免在工作流初始化阶段(规则外)去匹配动态生成的文件,这类操作要放到规则内部或依赖完成后执行。
内容的提问来源于stack exchange,提问作者QueNuevo
相关产品推荐
相关产品推荐

