为何Python多线程未提升程序运行速度?
问题原因分析
1. Python GIL(全局解释器锁)的限制
你的任务属于CPU密集型计算(大量求和、除法运算),而Python的线程受GIL约束:同一时间只有一个线程能执行Python字节码,多线程无法真正利用多核CPU并行计算,反而会因为线程切换带来额外开销,导致效率不升反降。
2. 伪并行:两个线程池是串行执行的
你代码里的两个ThreadPoolExecutor是先后运行的——第一个with块执行完毕后,才会启动第二个线程池。这意味着整个"并行"阶段实际上是串行跑了两组线程任务,完全没有同时利用多线程的优势。
3. 代码缩进错误导致的潜在问题
list1和list2的定义缩进错误,被放在了dissect2函数内部的return语句之后,这部分代码永远不会被执行,实际运行时会抛出NameError。你能运行成功应该是实际代码中缩进是正确的,但这个错误会导致逻辑异常。
4. 任务量完全一致,串行无切换开销
并行阶段和串行阶段的任务量完全相同(都是处理range(1,9999)的所有元素),串行执行没有线程切换的额外开销,所以耗时甚至比并行略短。
解决方案
1. 改用多进程(ProcessPoolExecutor)处理CPU密集型任务
多进程可以绕过GIL,真正利用多核CPU并行计算,把注释掉的ProcessPoolExecutor启用,替换ThreadPoolExecutor:
from concurrent.futures import ProcessPoolExecutor # 启用多进程 # from concurrent.futures import ThreadPoolExecutor import time values = range(1,9999) list1 = range(0, 99999999) list2 = range(0, 99999999) def dissect1(qq): adding = sum(list1[qq:qq+100000]) averaging = adding / len(list1[qq:qq+100000]) return adding,averaging def dissect2(qq): adding = sum(list2[qq:qq+100000]) averaging = adding / len(list2[qq:qq+100000]) return adding,averaging if __name__ == '__main__': tic = time.time() # 同时启动两个进程池并行执行 with ProcessPoolExecutor(max_workers=8) as exe1, ProcessPoolExecutor(max_workers=8) as exe2: result = exe1.map(dissect1, range(1,5000)) result2 = exe2.map(dissect2, range(5000,9999)) toc = time.time() print(f'Parallel execution, completed in {toc - tic} seconds') for v in values: a,b = dissect1(v) tac = time.time() print(f'Serial execution, completed in {tac - toc} seconds')
2. 合并任务,减少线程/进程池的创建开销
可以把dissect1和dissect2合并为一个函数,统一处理不同的列表,避免创建多个池:
def dissect(qq, target_list): adding = sum(target_list[qq:qq+100000]) averaging = adding / len(target_list[qq:qq+100000]) return adding, averaging # 调用时传入不同列表 with ProcessPoolExecutor(max_workers=16) as exe: tasks = [(qq, list1) for qq in range(1,5000)] + [(qq, list2) for qq in range(5000,9999)] results = exe.map(lambda x: dissect(*x), tasks)
3. 修复缩进错误
确保list1和list2的定义在函数外部,避免未定义的错误。
内容的提问来源于stack exchange,提问作者Matthias Ho
相关产品推荐
相关产品推荐

