Python循环与函数优化请求:蒙特卡洛模拟代码性能提升
Python蒙特卡洛模拟代码性能优化求助
问题背景
我正在优化一段蒙特卡洛(MC)模拟的Python代码,通过jackedCodeTimerPY工具分析了各模块耗时,目前存在以下性能瓶颈,恳请优化建议:
scores.append(...)耗时占比70%(184s/272s),最终需要计算各项平均值,如何提升结果收集效率?- 外层
for i in range(0,100000)循环重复执行同一函数,能否提升效率或并行化? match_simulation函数因循环依赖无法向量化/并行化,已将stats计算移到大循环外,还有其他优化点吗?resolveAction函数通过score[xxx] +=1更新字典,有没有更高效的实现方式?
另外,相同逻辑的Excel实现速度更快,说明当前Python版本还有优化空间。
执行时间统计
label min max mean total run count ----------------- ------ ------------ ------------- ---------- ----------- Total 272.84 272.84 272.84 272.84 1 Loop 0 0.25997 0.000271895 271.895 1000000 match_simulation 0 0.25797 8.817e-05 88.17 1000000 Action loop 0 0.25797 6.96475e-05 69.6475 1000000 resolveAction 0 0.0290003 2.77846e-06 41.6769 15000000 pullBack 0 0.00300026 6.83808e-07 10.2571 15000000 Shuffling 0 0.00199914 9.55838e-06 9.55838 1000000 Initial asignment 0 0.00299644 4.98002e-06 4.98002 1000000 scoreCalc 0 0.0149989 8.25524e-07 0.825524 1000000
相关代码
match_simulation函数
def match_simulation(teamHStat, teamAStat, stats, JTimer): JTimer.start('match_simulation') JTimer.start('Initial asignment') act_stats = stats score = {"H_wins" : 0, "H_draws" : 0, "H_looses" : 0, 'H_goals' : 0, 'A_goals' : 0 , "H_rolled" : 0, "A_rolled" : 0, "H_not_rolled" : 0, "A_not_rolled" : 0 , "H_l" : 0, "H_r" : 0, "H_c" : 0, "H_sp" : 0 , "H_lg" : 0, "H_rg" : 0, "H_cg" : 0, "H_spg" : 0 , "H_ca_l" : 0, "H_ca_r" : 0, "H_ca_c": 0, "H_ca_sp": 0 , "H_ca_lg" : 0, "H_ca_rg" : 0, "H_ca_cg": 0, "H_ca_spg": 0 , "A_l" : 0, "A_r" : 0, "A_c" : 0, "A_sp" : 0 , "A_lg" : 0, "A_rg" : 0, "A_cg" : 0, "A_spg" : 0 , "A_ca_l" : 0, "A_ca_r" : 0, "A_ca_c": 0, "A_ca_sp": 0 , "A_ca_lg" : 0, "A_ca_rg" : 0, "A_ca_cg": 0, "A_ca_spg": 0 , "H_common" : 0, "A_common" : 0, "H_own" : 0, "A_own" : 0 , 'H_miss' : 0, "A_miss" : 0, 'H_ca_miss' : 0, "A_ca_miss" : 0 , "H_PDIM": 0, "A_PDIM" : 0, "H_PNF" : 0, "A_PNF" : 0, "H_PNF_miss" : 0, "A_PNF_miss" : 0 , "H_ca_not_rolled" : 0, "A_ca_not_rolled" : 0, "H_ca_rolled" : 0, "A_ca_rolled" : 0} action_types = ['H', 'H', 'H', 'H', 'H', 'A', 'A', 'A', 'A', 'A', 'Common', 'Common', 'Common', 'Common', 'Common'] JTimer.stop('Initial asignment') JTimer.start('Shuffling') shuffle(action_types) JTimer.stop('Shuffling') JTimer.start('Action loop') for action_type in action_types: prev_score = score JTimer.start('resolveAction') score = resolveAction(act_stats, action_type, score) JTimer.stop('resolveAction') JTimer.start('pullBack') if(score['H_goals'] > prev_score['H_goals'] or score['A_goals'] > prev_score['A_goals']): act_stats = stats_calculation( pullBack(score['H_goals'], score['A_goals'], teamHStat) , pullBack(score['A_goals'], score['H_goals'], teamAStat)) JTimer.stop('pullBack') JTimer.stop('Action loop') JTimer.start('scoreCalc') if(score['H_goals'] > score['A_goals']): score['H_wins'] += 1 elif score['H_goals'] == score['A_goals']: score['H_draws'] += 1 else: score['H_looses'] += 1 JTimer.stop('scoreCalc') JTimer.stop('match_simulation') return score
主循环代码
JTimer = JackedTiming() JTimer.start('Total') scores = [] stats = stats_calculation(teamHstats.iloc[0], teamAstats.iloc[0]) for i in range(0,100000): JTimer.start('Loop') scores.append(match_simulation(teamHstats.iloc[0], teamAstats.iloc[0], stats, JTimer)) JTimer.stop('Loop') JTimer.stop('Total') print(JTimer.report())
优化建议
1. 优化结果收集(解决scores.append高耗时)
既然最终只需要各项平均值,完全不需要存储10万份完整的结果字典,直接实时累加统计值:
# 先获取初始score的所有键,初始化累加器 initial_score = {"H_wins" : 0, ...} # 和原来的score初始结构一致 total_scores = {k: 0 for k in initial_score.keys()} stats = stats_calculation(teamHstats.iloc[0], teamAstats.iloc[0]) for i in range(100000): result = match_simulation(teamHstats.iloc[0], teamAstats.iloc[0], stats, JTimer) # 直接累加每个统计项 for key in total_scores: total_scores[key] += result[key] # 最后计算平均值 avg_scores = {k: v / 100000 for k, v in total_scores.items()}
这种方式彻底避免了append操作和大量内存占用,能直接砍掉70%的耗时。如果必须保留中间结果,改用collections.deque代替列表,append效率更高。
2. 外层循环的并行化优化
外层10万次模拟是完全独立的,不受内部循环依赖影响,适合用多进程并行(CPU密集型任务避坑GIL):
from concurrent.futures import ProcessPoolExecutor # 封装模拟函数,适配多进程参数传递 def simulation_wrapper(args): teamHStat, teamAStat, stats = args # 去掉计时工具(多进程下计时会混乱,若需统计可在进程内单独处理) return match_simulation(teamHStat, teamAStat, stats, None) if __name__ == "__main__": stats = stats_calculation(teamHstats.iloc[0], teamAstats.iloc[0]) # 生成10万次模拟的参数列表 args_list = [(teamHstats.iloc[0], teamAstats.iloc[0], stats) for _ in range(100000)] # 启动进程池,进程数建议设为CPU核心数的1-2倍 with ProcessPoolExecutor(max_workers=4) as executor: results = list(executor.map(simulation_wrapper, args_list)) # 后续处理结果(比如累加计算平均值)
测试时可以先跑1000次验证效率,再放大到10万次。
3. match_simulation函数内部优化
- 复用
action_types列表:把列表定义提到函数外作为全局常量,每次模拟只复制并打乱,避免重复创建列表的开销:# 全局定义 ACTION_TYPES = ['H']*5 + ['A']*5 + ['Common']*5 def match_simulation(...): # ... action_types = ACTION_TYPES.copy() shuffle(action_types) # ... - 修改
resolveAction为原地更新:当前每次调用resolveAction都返回新字典,会产生大量内存分配开销,改成直接修改传入的score字典,不返回新对象。 - 简化进球判断逻辑:让
resolveAction返回是否产生进球的布尔值,不用每次对比prev_score和score的进球数,减少字典键查找次数。 - 移除冗余计时代码:性能分析完成后,把
JTimer的所有start/stop注释掉,减少函数调用开销。
4. resolveAction的字典更新优化
- 用
dataclass替代字典:属性访问速度远快于字典键查找,定义一个Score数据类来存储统计项:from dataclasses import dataclass @dataclass class Score: H_wins: int = 0 H_draws: int = 0 H_looses: int = 0 H_goals: int = 0 # ... 其他所有统计项,和原字典键对应 @classmethod def reset(cls): return cls() # 返回初始化全为0的实例 # 在match_simulation里初始化 score = Score.reset() # 在resolveAction里更新 score.H_wins += 1 - 减少键查找次数:如果
resolveAction里有多个连续更新同一类统计项的逻辑,尽量合并,避免重复查找键。
其他通用优化
- 切换到PyPy运行:PyPy的JIT编译器对循环密集型代码优化极强,这种蒙特卡洛模拟可能获得数倍速度提升。
- 内联小函数:如果
resolveAction逻辑不复杂,把它的代码内联到match_simulation的循环里,减少1500万次函数调用的开销。 - 优化
shuffle操作:可以预先生成多组打乱的action_types列表,循环时直接取用,避免每次模拟都调用shuffle。
内容的提问来源于stack exchange,提问作者Keru
相关产品推荐
相关产品推荐

