如何在Python中自定义冒号结尾控制语句?用于Snakemake开发
在Python中模拟Snakemake风格的自定义冒号语句
需求说明
想要在Python里实现类似Snakemake中input:、output:这类以冒号结尾的自定义语法:
input: table="table.csv" genome="Genome/genome.fa"
自动转换成对应的SM对象赋值:
input = SM(table="table.csv", genome="Genome/genome.fa")
目前已有SM类模拟Snakemake的相关对象,但需要手动编写赋值语句,希望简化为冒号加缩进的语法。
解决方案
Python本身不支持直接扩展语法规则,因此需要通过源代码预处理的方式,将自定义的冒号语句转换为合法的Python赋值代码。我们可以写一个装饰器来自动完成这个转换:
1. 优化SM类
首先简化原有的SM类,直接将关键字参数绑定到实例属性:
class SM: def __init__(self, **kwargs): # 把所有关键字参数直接设为实例属性 self.__dict__.update(kwargs)
2. 实现语法转换装饰器
这个装饰器会扫描函数内的源代码,将xxx:形式的语句加上后续缩进的键值对,转换为xxx = SM(...)的赋值语句:
import inspect def snakemake_syntax(func): # 获取函数的源代码 source_lines = inspect.getsource(func).split('\n') processed_lines = [] in_block = False current_var = None arg_list = [] for line in source_lines: stripped_line = line.strip() # 识别以冒号结尾的行(排除注释) if not in_block and stripped_line.endswith(':') and not stripped_line.startswith('#'): current_var = stripped_line[:-1].strip() in_block = True arg_list = [] continue # 处理缩进块内的键值对 if in_block: # 跳过空行和注释 if not stripped_line or stripped_line.startswith('#'): processed_lines.append(line) continue # 收集键值对(去掉末尾可能的逗号) arg_list.append(stripped_line.rstrip(',')) # 检查当前行是否是块的最后一行(非缩进行) if not line.startswith(' ') and stripped_line: # 生成赋值语句 processed_lines.append(f"{current_var} = SM({', '.join(arg_list)})") in_block = False current_var = None continue # 处理普通行,同时检查是否有未闭合的块 if current_var is not None: processed_lines.append(f"{current_var} = SM({', '.join(arg_list)})") current_var = None processed_lines.append(line) # 处理函数末尾未闭合的块 if in_block and current_var is not None: processed_lines.append(f"{current_var} = SM({', '.join(arg_list)})") # 执行处理后的代码,更新函数的局部变量 exec('\n'.join(processed_lines), func.__globals__, func.__locals__) return func
3. 使用示例
用装饰器包裹你的规则函数,就可以直接使用Snakemake风格的冒号语法:
@snakemake_syntax def process_genome(): input: table="rmats/binding_strength.maxent.CLIP.csv" genome="Genome/genome.fa" output: filtered_table="output/filtered_table.csv" # 直接使用input和output的属性,无需修改即可迁移到Snakefile import pandas as pd df = pd.read_csv(input.table, index_col=0) # 处理数据... df.to_csv(output.filtered_table) process_genome()
原理说明
Snakemake本身也是通过自定义解析器将Snakefile转换为Python代码执行的,我们这里的装饰器做了类似的工作:扫描源代码中的自定义语法,将其转换为合法的Python赋值语句,再执行转换后的代码。这种方法可以让你在Python脚本中使用和Snakefile一致的语法,方便代码迁移。
内容的提问来源于stack exchange,提问作者user3392394
相关产品推荐
相关产品推荐

