You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snakemake通配符约束:如何排除以_pe结尾的样本名

在Snakemake中定义排除特定后缀的通配符约束

方法1:通过wildcard_constraints使用负向正则断言

Snakemake支持在规则中通过wildcard_constraints参数为通配符指定正则约束,负向前瞻/后顾属于Python正则的特性,只要写法正确就能生效。针对你的需求,要匹配不以_pe结尾的样本名,可使用负向后瞻正则:

rule single_end_align:
    input:
        "raw_data/{sample_sp}.fastq"
    output:
        "aligned/{sample_sp}.bam"
    wildcard_constraints:
        # 负向后瞻:确保字符串结尾不是_pe
        sample_sp=r'^.*(?<!_pe)$'
    shell:
        """
        # 单端数据比对命令示例
        bowtie2 -x ref_genome -U {input} | samtools view -bS - > {output}
        """

验证正则有效性

可先在Python中测试正则是否符合预期:

import re
pattern = re.compile(r'^.*(?<!_pe)$')
print(pattern.match("sample1"))       # 匹配成功(单端样本)
print(pattern.match("sample2_pe"))    # 匹配失败(双端样本)

若正则测试通过但Snakemake仍无法匹配,检查文件路径是否完全对应(比如单端文件是否确实是raw_data/{sample_sp}.fastq格式,无额外后缀)。


方法2:预筛选样本列表(更直观可靠)

如果正则约束的方式容易出问题,推荐提前通过Python代码筛选单端/双端样本列表,再用expand函数生成规则的输入输出,这种方式逻辑更清晰,也避免了正则匹配的不确定性:

# 在Snakefile开头添加样本筛选逻辑
import glob

# 筛选单端样本:文件名(去掉.fastq后缀)不以_pe结尾
single_end_samples = [
    fname.split("/")[-1].rstrip(".fastq") 
    for fname in glob.glob("raw_data/*.fastq")
    if not fname.split("/")[-1].rstrip(".fastq").endswith("_pe")
]

# 筛选双端样本:匹配*_pe_1.fastq格式的文件,提取样本名
paired_end_samples = [
    fname.split("/")[-1].replace("_1.fastq", "") 
    for fname in glob.glob("raw_data/*_pe_1.fastq")
]

# 总规则
rule all:
    input:
        expand("aligned/{sample}.bam", sample=single_end_samples),
        expand("aligned/{sample}.bam", sample=paired_end_samples)

# 单端分析规则
rule single_end_align:
    input:
        "raw_data/{sample}.fastq"
    output:
        "aligned/{sample}.bam"
    shell:
        """
        bowtie2 -x ref_genome -U {input} | samtools view -bS - > {output}
        """

# 双端分析规则
rule paired_end_align:
    input:
        "raw_data/{sample}_1.fastq",
        "raw_data/{sample}_2.fastq"
    output:
        "aligned/{sample}.bam"
    shell:
        """
        bowtie2 -x ref_genome -1 {input[0]} -2 {input[1]} | samtools view -bS - > {output}
        """

这种方式直接通过代码明确区分样本类型,Snakemake仅处理预定义的样本列表,调试和维护都更简单。


为什么负向正则可能不生效?

Snakemake的通配符匹配需要对应到文件路径中的具体字符串片段,若正则仅包含断言(无实际匹配内容),或者文件路径结构与正则预期不符,可能导致匹配失败。比如如果单端文件路径包含子目录,正则需要调整以适配完整的通配符匹配范围。

内容的提问来源于stack exchange,提问作者Whitehot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 07:47:24