Snakemake通配符约束:如何排除以_pe结尾的样本名
在Snakemake中定义排除特定后缀的通配符约束
方法1:通过wildcard_constraints使用负向正则断言
Snakemake支持在规则中通过wildcard_constraints参数为通配符指定正则约束,负向前瞻/后顾属于Python正则的特性,只要写法正确就能生效。针对你的需求,要匹配不以_pe结尾的样本名,可使用负向后瞻正则:
rule single_end_align: input: "raw_data/{sample_sp}.fastq" output: "aligned/{sample_sp}.bam" wildcard_constraints: # 负向后瞻:确保字符串结尾不是_pe sample_sp=r'^.*(?<!_pe)$' shell: """ # 单端数据比对命令示例 bowtie2 -x ref_genome -U {input} | samtools view -bS - > {output} """
验证正则有效性
可先在Python中测试正则是否符合预期:
import re pattern = re.compile(r'^.*(?<!_pe)$') print(pattern.match("sample1")) # 匹配成功(单端样本) print(pattern.match("sample2_pe")) # 匹配失败(双端样本)
若正则测试通过但Snakemake仍无法匹配,检查文件路径是否完全对应(比如单端文件是否确实是raw_data/{sample_sp}.fastq格式,无额外后缀)。
方法2:预筛选样本列表(更直观可靠)
如果正则约束的方式容易出问题,推荐提前通过Python代码筛选单端/双端样本列表,再用expand函数生成规则的输入输出,这种方式逻辑更清晰,也避免了正则匹配的不确定性:
# 在Snakefile开头添加样本筛选逻辑 import glob # 筛选单端样本:文件名(去掉.fastq后缀)不以_pe结尾 single_end_samples = [ fname.split("/")[-1].rstrip(".fastq") for fname in glob.glob("raw_data/*.fastq") if not fname.split("/")[-1].rstrip(".fastq").endswith("_pe") ] # 筛选双端样本:匹配*_pe_1.fastq格式的文件,提取样本名 paired_end_samples = [ fname.split("/")[-1].replace("_1.fastq", "") for fname in glob.glob("raw_data/*_pe_1.fastq") ] # 总规则 rule all: input: expand("aligned/{sample}.bam", sample=single_end_samples), expand("aligned/{sample}.bam", sample=paired_end_samples) # 单端分析规则 rule single_end_align: input: "raw_data/{sample}.fastq" output: "aligned/{sample}.bam" shell: """ bowtie2 -x ref_genome -U {input} | samtools view -bS - > {output} """ # 双端分析规则 rule paired_end_align: input: "raw_data/{sample}_1.fastq", "raw_data/{sample}_2.fastq" output: "aligned/{sample}.bam" shell: """ bowtie2 -x ref_genome -1 {input[0]} -2 {input[1]} | samtools view -bS - > {output} """
这种方式直接通过代码明确区分样本类型,Snakemake仅处理预定义的样本列表,调试和维护都更简单。
为什么负向正则可能不生效?
Snakemake的通配符匹配需要对应到文件路径中的具体字符串片段,若正则仅包含断言(无实际匹配内容),或者文件路径结构与正则预期不符,可能导致匹配失败。比如如果单端文件路径包含子目录,正则需要调整以适配完整的通配符匹配范围。
内容的提问来源于stack exchange,提问作者Whitehot
相关产品推荐
相关产品推荐

