优化运行缓慢的Python Genotype对比函数求助
Python基因型对比函数优化求助
我有一个运行速度极慢的Python函数,用于对比两个Genotype对象并返回样本基因型字典,其运行时长受Genotype对象大小及sample_list长度影响。以下是该函数的当前实现:
def compare_gt(gt1: Genotype, gt2: Genotype, sample_list: list = None): # is gt1 is None if not gt1 is None and not gt1.isValid(): gt1=None if not gt2 is None and not gt2.isValid(): gt2=None if gt1 is None and gt2 is None: return None if gt1 is None: if sample_list is None: sample_list = gt2.get_sample_list() return {sample: "6" + gt2.get_sample_genotypes(sample_list)[sample] for sample in sample_list} if gt2 is None: if sample_list is None: sample_list = gt1.get_sample_list() return {sample: gt1.get_sample_genotypes(sample_list)[sample] + "6" for sample in sample_list} # compare two genotype if sample_list is None: sample_list = gt1.get_sample_list() return {sample: gt1.get_sample_genotypes(sample_list)[sample] + gt2.get_sample_genotypes(sample_list)[sample] for sample in sample_list}
我已尝试部分优化手段,比如避免重复计算sample_list、将gt1和gt2的结果存入变量以避免重复调用方法,但函数运行仍十分缓慢,希望能得到进一步优化建议。
更新内容
根据建议补充程序运行逻辑伪代码:
class gt_reader(): def __init__(self, ...): # do something def update(self): # read next genotype line def get_sample_genotypes(self, sample_list) -> dict: # do something return sample_gt # other methods gt1 = gt_reader(...) gt2 = gt_reader(...) while gt1.update(): # some code while gt2.update(): # some code compare_gt(gt1, gt2) # some code if x: break
除了待获取的优化建议外,我还对程序做了额外优化:确保get_sample_genotypes()的输出始终按查询排序,当gt1和gt2样本相同时无需持续通过索引获取返回值,这在大量对比操作中已提升了部分性能。
内容的提问来源于stack exchange,提问作者zhang
相关产品推荐
相关产品推荐

