You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何简化Snakemake中重复的样本表查询输入函数?

简化Snakemake样本表查询输入函数的方法

针对你在Snakemake中重复编写样本表查询输入函数的问题,最简洁的方法是用**闭包(嵌套函数)**或者functools.partial来批量生成针对不同列的查询函数,既能避免重复代码,又能满足命名input的要求。

方法1:闭包生成专用查询函数

先定义一个接受目标列名的外层函数,返回一个只接受wildcards参数的内层函数,这样就能直接用于命名input项:

import pandas as pd

# 假设你的样本表已经加载为DataFrame
sample_table = pd.read_csv("samples.csv")

def make_sample_lookup(target_col):
    def lookup(wildcards):
        # 根据wildcards中的样本ID匹配行,返回目标列的值
        return sample_table.loc[sample_table["sample_id"] == wildcards.sample, target_col].iloc[0]
    return lookup

# 生成针对不同列的查询函数
get_image = make_sample_lookup("image_path")
get_visium_fastqs = make_sample_lookup("fastq_paths")
get_sample_barcode = make_sample_lookup("barcode_file")

在规则中直接使用这些生成的函数,不管是命名还是非命名input都能正常工作:

rule process_visium:
    input:
        image=get_image,
        fastqs=get_visium_fastqs,
        barcode=get_sample_barcode
    output:
        "processed/{sample}.h5ad"
    shell:
        """
        # 处理命令
        """

方法2:用functools.partial绑定参数

如果更喜欢用偏函数的方式,可以借助functools.partial固定目标列参数,生成符合Snakemake要求的单参数函数:

from functools import partial
import pandas as pd

sample_table = pd.read_csv("samples.csv")

def lookup_sample_table(wildcards, target_col):
    return sample_table.loc[sample_table["sample_id"] == wildcards.sample, target_col].iloc[0]

# 绑定目标列,生成专用函数
get_image = partial(lookup_sample_table, target_col="image_path")
get_visium_fastqs = partial(lookup_sample_table, target_col="fastq_paths")

规则使用方式和闭包方法完全一致,同样支持命名input。

关键优势

  • 只需要写一次核心查询逻辑,避免重复代码
  • 生成的函数命名清晰,可读性强
  • 完全兼容Snakemake的命名/非命名input要求

内容的提问来源于stack exchange,提问作者j0hn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 02:57:54