如何加速基于列表统计的共识判断Python代码?
优化共识判断函数的性能建议
我正在优化一个共识判断函数,它接收字符串列表(比如['bull','bull','bear'])作为输入,需要得出“多数共识”,输出字符串'bull'。想请教如何提升它的运行速度。
原代码
consensus_type = "majority" activeIndicators = ["bull","bull","bear"] def consensus( activeIndicators ): counterBull = activeIndicators.count("bull") counterBear = activeIndicators.count("bear") counterNeutral = activeIndicators.count("neutral") lists = [counterBull,counterBear,counterNeutral] max_value_from_list = max(lists) count_of_max_value_in_list = lists.count(max_value_from_list) if consensus_type == "majority": if count_of_max_value_in_list == 1: d = {'bull':counterBull,'bear':counterBear,'neutral':counterNeutral} consensus = max(d, key=d.get) else: consensus = "neutral" elif consensus_type == "unanimity": if max_value_from_list == len(activeIndicators): d = {'bull':counterBull,'bear':counterBear,'neutral':counterNeutral} consensus = max(d, key=d.get) else: consensus = "neutral" return consensus
已完成的优化(速度提升至原版本两倍)
以下是简化后的优化代码,目前运行速度是原版本的两倍,希望得到更多优化建议:
def consensus( activeIndicators ): counterBull = activeIndicators.count("bull") counterBear = activeIndicators.count("bear") counterNeutral = len(activeIndicators) - counterBull - counterBear consensus = "neutral" if counterBull >= counterBear and counterBull >= counterNeutral: if consensus_type == 'unanimity' and counterBull == len(activeIndicators): consensus = "bull" elif consensus_type == 'majority' and (counterBull != counterBear and counterBull != counterNeutral): consensus = "bull" elif counterBear >= counterBull and counterBear >= counterNeutral: if consensus_type == 'unanimity' and counterBear == len(activeIndicators): consensus = "bear" elif consensus_type == 'majority' and (counterBear != counterBull and counterBear != counterNeutral): consensus = "bear" return consensus
进一步优化建议
1. 单次遍历完成计数,减少列表遍历次数
原代码和当前优化代码都调用了两次count()方法,每次count()都会完整遍历列表一次。改成单次遍历统计所有类型数量,能直接减少一半的遍历开销,列表越长收益越明显:
def consensus(activeIndicators, consensus_type="majority"): counterBull = 0 counterBear = 0 for item in activeIndicators: if item == "bull": counterBull += 1 elif item == "bear": counterBear += 1 counterNeutral = len(activeIndicators) - counterBull - counterBear consensus = "neutral" if counterBull >= counterBear and counterBull >= counterNeutral: if consensus_type == 'unanimity' and counterBull == len(activeIndicators): consensus = "bull" elif consensus_type == 'majority' and (counterBull != counterBear and counterBull != counterNeutral): consensus = "bull" elif counterBear >= counterBull and counterBear >= counterNeutral: if consensus_type == 'unanimity' and counterBear == len(activeIndicators): consensus = "bear" elif consensus_type == 'majority' and (counterBear != counterBull and counterBear != counterNeutral): consensus = "bear" return consensus
2. 将consensus_type作为函数参数传入
当前代码中consensus_type是全局变量,既降低了函数的封装性,全局变量的访问速度也略慢于局部变量。把它作为参数传入,不仅让函数更易复用,还能小幅提升性能。
3. 提前按consensus_type分支,减少嵌套层级
调整判断逻辑顺序,先根据consensus_type分支,再处理计数后的判断,既能让代码逻辑更清晰,也能减少运行时的嵌套条件判断开销:
def consensus(activeIndicators, consensus_type="majority"): counterBull = 0 counterBear = 0 for item in activeIndicators: if item == "bull": counterBull += 1 elif item == "bear": counterBear += 1 counterNeutral = len(activeIndicators) - counterBull - counterBear if consensus_type == 'unanimity': if counterBull == len(activeIndicators): return "bull" elif counterBear == len(activeIndicators): return "bear" return "neutral" elif consensus_type == 'majority': if counterBull > counterBear and counterBull > counterNeutral: return "bull" elif counterBear > counterBull and counterBear > counterNeutral: return "bear" return "neutral"
4. 处理长列表时使用collections.Counter
如果需要处理的列表规模很大,使用标准库的collections.Counter会更高效——它的底层实现是C语言,统计速度比纯Python遍历更快:
from collections import Counter def consensus(activeIndicators, consensus_type="majority"): cnt = Counter(activeIndicators) counterBull = cnt.get("bull", 0) counterBear = cnt.get("bear", 0) counterNeutral = cnt.get("neutral", 0) if consensus_type == 'unanimity': if counterBull == len(activeIndicators): return "bull" elif counterBear == len(activeIndicators): return "bear" return "neutral" elif consensus_type == 'majority': if counterBull > counterBear and counterBull > counterNeutral: return "bull" elif counterBear > counterBull and counterBear > counterNeutral: return "bear" return "neutral"
注意:如果列表很短,引入Counter的初始化开销可能超过收益,因此更适合大规模数据场景。
内容的提问来源于stack exchange,提问作者zolp
相关产品推荐
相关产品推荐

