Snakemake访问样本表格元数据出现IndexError报错如何解决
问题解决方案
核心错误原因
expand函数的样本列表取值错误:直接遍历Pandas DataFrame对象默认返回的是表的列名列表,你的表格列名为Sample和Layout,因此原代码实际生成的目标文件是Sample.txt和Layout.txt,而非你需要的样本名对应的txt。当Snakemake匹配到Sample.txt目标时,通配符sample取值为Sample,元表中没有对应样本行,取.Layout[0]时就触发了索引越界错误。- 取值逻辑冗余:你已经将
Sample列设为DataFrame的索引,不需要额外写布尔匹配逻辑取属性,直接用索引查询更稳定。
修正后的完整Snakefile代码
configfile: "config.yaml" import pandas as pd sample_file = config["sample_file"] samples = pd.read_table(sample_file).set_index("Sample", drop = True) rule all: input: # 取DataFrame的索引(即样本名列表)传给expand expand("{sample}.txt", sample=samples.index) rule rule1: output: "{sample}.txt" params: # 直接用索引查询对应Layout属性,逻辑更简洁 tag = lambda wc: samples.loc[wc.sample, "Layout"] shell: """ echo {params.tag} > {output} """
小优化:
echo重定向本身就会自动创建输出文件,不需要额外写touch {output}语句。
运行验证
执行snakemake -c1后,会生成SRR11213896.txt和ERR3887380.txt两个文件,内容分别为SE和PE,完全符合预期。
内容的提问来源于stack exchange,提问作者Hannah
相关产品推荐
相关产品推荐

