如何修复Multiprocessing Pool导致CPU占用100%的问题
批量网站扫描脚本CPU占用100%问题分析与修复方案
我已花费数日仍无法解决Python脚本的问题。该脚本仅用于在批量网站(数量范围2000-50000,测试200个网站时问题依旧)上搜索特定文件,但运行时CPU占用率始终达100%。以下是脚本中与Multiprocessing相关的核心代码片段:
try: import os import easygui import pyautogui as py import datetime import pwinput import json from queue import Queue from queue import Empty from collections import Counter import random import string import threading import subprocess import multiprocessing from multiprocessing import Process, Pool, Value, Lock import grequests import requests from requests.exceptions import ConnectionError from requests.exceptions import HTTPError import time from os import path from zipfile import * from pathlib import Path from time import sleep import socket from datetime import datetime import re from re import findall as reg import urllib.request as urlrq from urllib.parse import urlparse import certifi import ssl import urllib3 from urllib3.exceptions import InsecureRequestWarning from urllib3 import disable_warnings disable_warnings(InsecureRequestWarning) import validators from apscheduler.schedulers.background import BackgroundScheduler import colorama from colorama import Fore, Back, Style, init import itertools import ctypes except ImportError: import os os.system('pip install grequests') os.system('pip install pwinput') os.system('pip install colorama') os.system('pip install apscheduler') os.system('pip install requests') os.system('pip install findall') os.system('pip install http.client') os.system('pip install psutil') os.system('pip install validators') os.system('pip install itertools') os.system('pip install urllib3') os.system('pip install easygui') os.system('pip install pyautogui') os.system('pip install zipfile') colorama.init(autoreset=True) R = Fore.RED P = Fore.MAGENTA A = Fore.CYAN G = Fore.GREEN B = Fore.BLUE W = Fore.WHITE Y = Fore.YELLOW C = Fore.RESET CC = Style.RESET_ALL PP = Style.BRIGHT + Fore.LIGHTMAGENTA_EX AA = Style.BRIGHT + Fore.LIGHTCYAN_EX RR = Style.BRIGHT + Fore.LIGHTRED_EX YY = Style.BRIGHT + Fore.LIGHTYELLOW_EX lock = threading.Lock() NUM_PROCESSES = int() class Counter(object): def __init__(self, initval=0): self.val = Value('i', initval) self.lock = Lock() def increment(self): with self.lock: self.val.value += 1 def value(self): with self.lock: return self.val.value def updateTitle(NUM_PROCESSES,counterhits,counterdone,countersl,counterml,counter_listok,username): while True: hits = int(counterhits.value()) done = int(counterdone.value()) shtot = int(countersl.value()) maitot = int(counterml.value()) listok = num_listok + int(counter_listok) remain_scan = listok - hits elapsed = time.strftime('%H:%M:%S', time.gmtime(time.time() - start)) ctypes.windll.kernel32.SetConsoleTitleW(f'Site Valid For: {listok} | Started: {hits} | Complete: {done} | Remain: {remain_scan} | Sl Found: {shtot} | Ml Found: {maitot} | Threads: {NUM_PROCESSES} | Time elapsed: {elapsed}') sleep(0.3) def worker_main(queue,counterhits,counterdone,countersl,counterml): while True: try: site = queue.get(block=False) except queue.Empty: sleep(0.5) continue except queue.Full: sleep(0.5) continue if site is None: break start_time_base = time.process_time() counterhits.increment() do_work(site,countersl,counterml,counterdone) counterdone.increment() end_time_base = time.process_time() start = time.time() def main(): global username listok = int(len(filter_data)) the_queue = multiprocessing.Queue() counterhits = Counter(0) counterdone = Counter(0) countersl = Counter(0) counterml = Counter(0) counter_listok = int(listok) the_pool = multiprocessing.Pool(NUM_PROCESSES, worker_main,(the_queue,counterhits,counterdone,countersl,counterml)) procs = [Process(target=updateTitle, args=(NUM_PROCESSES,counterhits,counterdone,countersl,counterml,counter_listok,username), daemon=True) for i in range(1)] prefix = ['http://'] for i in range(NUM_PROCESSES): for site_il in filter_data: site_or = site_il.rstrip("\n") if (site_or.startswith("http://")) : site_or = site_or.replace("http://","") elif (site_or.startswith("https://")) : site_or = site_or.replace("https://","") site_or = site_or.rstrip() site_or = site_or.split('/')[0] if ('www.' in site_or) : site_or = site_or.replace("www.", "") sitexx = [sub + site_or for sub in prefix] for site in sitexx: the_queue.put(site) the_queue.put(None) for p in procs: p.start() the_queue.close() the_queue.join_thread() the_pool.close() the_pool.join() end_time_base = time.perf_counter() for p in procs: p.join() os.system('pause>nul')
问题根源分析
- 未初始化的进程数参数:
NUM_PROCESSES = int()会将进程数设为0,此时multiprocessing.Pool默认使用CPU核心数,但后续任务填充逻辑错误,导致进程无意义空转。 - 进程池初始化逻辑错误:将
worker_main作为Pool的初始化函数传入,导致进程池的每个子进程启动后直接进入死循环,完全违背进程池的设计意图。 - 任务队列重复填充:外层
for i in range(NUM_PROCESSES):循环会将所有网站任务重复添加NUM_PROCESSES次,任务量暴增的同时,每个进程都收到终止信号None,导致逻辑混乱,持续占用CPU。 - Worker的忙等待轮询:
queue.get(block=False)在队列空时触发异常后立即重试,这种轮询方式让CPU在空闲时持续空转,直接拉满占用率。 - 冗余模块与锁竞争:导入大量未使用的模块(如
pyautogui、easygui),自定义Counter类的锁逻辑存在潜在竞争,进一步加剧CPU开销。
修复方案
1. 合理设置进程/线程数
网络IO密集型任务无需过多CPU核心,建议设置为CPU核心数的2-4倍,或改用开销更低的线程池:
# 进程数设置示例 NUM_PROCESSES = multiprocessing.cpu_count() * 3 # 或改用线程池(更适合IO密集任务) from concurrent.futures import ThreadPoolExecutor
2. 重构进程池/线程池使用逻辑
放弃“进程池+队列”的错误组合,改用线程池直接映射任务:
def process_site(site, countersl, counterml, counterdone): counterhits.increment() do_work(site, countersl, counterml, counterdone) counterdone.increment() def main(): # ... 其他初始化逻辑 ... sitexx_list = [] # 提前生成所有处理后的site列表 for site_il in filter_data: site_or = site_il.rstrip("\n") if site_or.startswith(("http://", "https://")): site_or = site_or.split("://")[1] site_or = site_or.rstrip().split('/')[0].replace("www.", "") sitexx_list.append(f"http://{site_or}") # 用线程池批量处理任务 with ThreadPoolExecutor(max_workers=NUM_PROCESSES) as executor: for site in sitexx_list: executor.submit(process_site, site, countersl, counterml, counterdone)
3. 修正任务填充逻辑(保留队列方案时)
去掉重复填充队列的外层循环,仅生成一次任务列表,再添加对应数量的终止信号:
# 生成所有处理后的site sitexx_list = [] for site_il in filter_data: site_or = site_il.rstrip("\n") if site_or.startswith(("http://", "https://")): site_or = site_or.split("://")[1] site_or = site_or.rstrip().split('/')[0].replace("www.", "") sitexx_list.append(f"http://{site_or}") # 填充队列 for site in sitexx_list: the_queue.put(site) # 给每个Worker添加终止信号 for _ in range(NUM_PROCESSES): the_queue.put(None)
4. 修改Worker的队列获取方式
用block=True替代非阻塞获取,让进程在队列空时阻塞,避免轮询空转:
def worker_main(queue,counterhits,counterdone,countersl,counterml): while True: try: site = queue.get(block=True) # 阻塞等待任务,无空转 except queue.Full: sleep(0.5) continue if site is None: break counterhits.increment() do_work(site,countersl,counterml,counterdone) counterdone.increment()
5. 清理冗余模块与优化计数器
- 删除未使用的模块导入(如
pyautogui、easygui、apscheduler等) - 改用
multiprocessing.Manager提供的内置计数器,简化锁逻辑:
from multiprocessing import Manager def main(): manager = Manager() counterhits = manager.Counter() counterdone = manager.Counter() countersl = manager.Counter() counterml = manager.Counter()
内容的提问来源于stack exchange,提问作者Massimo B.
相关产品推荐
相关产品推荐

