Snakemake调用Python的seaborn生成多输出图时缺失输出如何解决
Snakemake绘图输出缺失问题修复方案
1. 核心报错原因分析
输出缺失由多个错误叠加导致:
- 你用
shell指令直接调用Python脚本,脚本中没有获取snakemake对象的合法渠道,直接调用snakemake.input/snakemake.output会直接抛出未定义错误,脚本执行失败自然不会生成输出 - Python脚本里的输入索引写错:
snakemake.input[O]是大写字母O,不是数字0,会直接报变量未定义 - 第二个输入是tsv格式,脚本里错误使用逗号作为分隔符,会导致读取文件失败
- 循环遍历列存储图片时,所有列的图都会覆盖同一个
snakemake.output[0]路径,最终只会保留最后一列的图,如果你预期每列对应一张图,这个逻辑完全错误 - 给出的Snakemake伪代码中input列表的两个文件路径没有加逗号分隔,本身就存在语法错误
2. 修复方案
2.1 输出路径的正确写法
两种合法方案二选一即可:
方案A:用Snakemake的script指令替代shell(推荐)
Snakemake的script指令会自动给Python脚本注入snakemake全局对象,不需要额外传参:
# 修正后的Snakemake规则片段 rule seaborn: input: "/OtherPath/to/file1.csv", "/AnotherPath/to/file2.tsv" output: "/path/to/outputs{wc1}_{wc2}_{col}.png" script: "SeabornPlot.py"
对应Python脚本修改为:
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns # 输入索引改为数字0,第二个输入的分隔符改为\t CSV1 = pd.read_csv(snakemake.input[0], delimiter=",") CSV2 = pd.read_csv(snakemake.input[1], delimiter="\t") # 从通配符获取当前要绘制的列名,避免循环覆盖 target_col = snakemake.wildcards.col g = sns.barplot(x=target_col, data=CSV2) fig = g.get_figure() fig.savefig(snakemake.output[0]) plt.close()
方案B:保留shell指令,通过参数传递路径
如果要继续用shell调用Python,需要把路径作为参数传入:
# 修正后的Snakemake规则片段 rule seaborn: input: "/OtherPath/to/file1.csv", "/AnotherPath/to/file2.tsv" output: "/path/to/outputs{wc1}_{wc2}_{col}.png" shell: "python SeabornPlot.py {input[0]} {input[1]} {output[0]} {wildcards.col}"
对应Python脚本修改为:
import sys import pandas as pd import matplotlib.pyplot as plt import seaborn as sns CSV1 = pd.read_csv(sys.argv[1], delimiter=",") CSV2 = pd.read_csv(sys.argv[2], delimiter="\t") output_path = sys.argv[3] target_col = sys.argv[4] g = sns.barplot(x=target_col, data=CSV2) fig = g.get_figure() fig.savefig(output_path) plt.close()
2.2 checkpoint使用合理性判断
你当前场景不需要用checkpoint,替换为普通rule即可。checkpoint仅适用于执行前无法确定输出数量/路径,需要执行完该步骤才能确定后续依赖的动态工作流场景,你当前的输出路径可以通过通配符提前确定,用checkpoint属于滥用,反而会增加工作流复杂度。
2.3 rule all的修改建议
如果你需要为CSV1的每一列都生成对应wc1、wc2的图,需要提前读取CSV1列名作为通配符取值,修改expand逻辑:
# 提前读取CSV1的列名 ALL_COLS = pd.read_csv("/OtherPath/to/file1.csv").columns.tolist() rule all: expand("/path/to/outputs{wc1}_{wc2}_{col}.png", wc1=WC1, wc2=WC2, col=ALL_COLS)
内容的提问来源于stack exchange,提问作者c00kieRaptor
相关产品推荐
相关产品推荐

