Python 3.5多线程实现是否正确?质数查找性能无提升疑问
嘿,你的多线程代码确实有问题,而且还得考虑Python GIL的限制!
首先看你创建线程的那两行代码——这是个新手常踩的坑:
thread1 = threading.Thread(target=findPrimes(testValues1)) thread2 = threading.Thread(target=findPrimes(testValues2))
你这里直接调用了findPrimes(testValues1),也就是说在创建线程之前,这两个函数已经同步执行完了!线程对象其实根本没机会干活,所以耗时和单线程完全一样,这是最直接的原因。
正确的写法应该是把函数本身传给target,然后用args参数传递列表,另外还要加join()让主线程等子线程执行完再计算时间:
thread1 = threading.Thread(target=findPrimes, args=(testValues1,)) thread2 = threading.Thread(target=findPrimes, args=(testValues2,)) thread1.start() thread2.start() thread1.join() # 等待线程1执行完毕 thread2.join() # 等待线程2执行完毕
但修正后你可能还是看不到提速——因为Python的GIL
即使你改对了线程的写法,在这个质数计算的CPU密集型任务里,多线程还是快不起来。这是因为CPython(我们日常用的Python解释器)有个全局解释器锁(GIL):同一时间只能有一个线程执行Python字节码。对于纯计算的任务,多线程其实是在单个CPU核心上交替跑,没法真正并行,所以速度没变化。
怎么用质数查找演示多线程的优势?
要体现多线程的价值,得换个场景,或者换个方式:
1. 改成IO密集型的质数查找
比如给每个质数检查加个模拟IO(比如等待网络响应、读写文件),这样线程在等IO的时候会释放GIL,让其他线程能干活:
import time import threading from math import sqrt from random import randint def findPrimesWithIO(aList): for testValue in aList: # 模拟IO等待,比如调用API的延迟 time.sleep(0.001) isPrime = True for i in range(2, int(sqrt(testValue)+1)): if testValue % i == 0: isPrime = False break # 找到因数就提前退出,优化下效率 if isPrime: pass # 生成测试数据 testValues1 = [randint(10, 100) for _ in range(1000)] testValues2 = [randint(10, 100) for _ in range(1000)] # 单线程跑 t = time.process_time() findPrimesWithIO(testValues1) findPrimesWithIO(testValues2) print('单线程耗时', time.process_time() - t) # 多线程跑 t = time.process_time() thread1 = threading.Thread(target=findPrimesWithIO, args=(testValues1,)) thread2 = threading.Thread(target=findPrimesWithIO, args=(testValues2,)) thread1.start() thread2.start() thread1.join() thread2.join() print('多线程耗时', time.process_time() - t)
这个版本里,多线程的耗时会明显比单线程少,因为一个线程等IO的时候,另一个线程可以继续计算。
2. 用多进程处理CPU密集型的质数计算
如果还是想做纯计算的质数查找,那就用multiprocessing模块——每个进程有自己的GIL,能真正利用多核CPU并行计算:
import time import multiprocessing from math import sqrt from random import randint def findPrimes(aList): primes = [] for testValue in aList: isPrime = True for i in range(2, int(sqrt(testValue)+1)): if testValue % i == 0: isPrime = False break if isPrime: primes.append(testValue) return primes # 生成更大的测试数据,不然差异不明显 testValues1 = [randint(100000, 1000000) for _ in range(1000)] testValues2 = [randint(100000, 1000000) for _ in range(1000)] # 单进程跑 t = time.process_time() findPrimes(testValues1) findPrimes(testValues2) print('单进程耗时', time.process_time() - t) # 多进程跑 t = time.process_time() p1 = multiprocessing.Process(target=findPrimes, args=(testValues1,)) p2 = multiprocessing.Process(target=findPrimes, args=(testValues2,)) p1.start() p2.start() p1.join() p2.join() print('多进程耗时', time.process_time() - t)
这个版本里,多进程会明显比单进程快,因为它用了多个CPU核心同时计算。
最后总结下你的问题
- 最开始的代码错误:创建线程时直接调用了函数,导致任务在主线程同步执行完了,线程根本没干活
- 即使修正线程写法,CPU密集型任务在CPython多线程下受GIL限制,无法并行提速
- 要演示多线程优势,要么改成IO密集型场景,要么用多进程处理CPU密集型任务
内容的提问来源于stack exchange,提问作者Andrew H
相关产品推荐
相关产品推荐

