Python多线程优化问询:如何高效批量检测存活域名/子域名?
问题描述
我编写了一个程序,用于检测subdomains.txt文件中的域名/子域名是否可通过HTTP/HTTPS协议存活。实现逻辑是向目标发送HEAD请求,若能获取响应状态码则判定为存活(对应load_url_http函数)。
为提升运行速度,我使用了concurrent.futures.ThreadPoolExecutor,初始设置线程数为200,但将线程数提升至300、400后,程序速度并未得到明显提升。我希望对程序进行优化,使其能够同时发送数千次请求。以下是我的程序源码及运行测试结果:
程序源码(python-request-multil.py)
import time import requests import concurrent.futures def load_url_http(protocol: str, domain: str, timeout: int = 10): try: conn = requests.head(protocol + "://" + domain, timeout=timeout) return conn.status_code except Exception: return None #--- main ---# start_time = time.time() worker = 400 protocol = "http" timeout = 10 print("Number of worker:", worker) with concurrent.futures.ThreadPoolExecutor(max_workers=worker) as executor: # The file object that the subdomain lives on will be written to file_live_subdomain = open("live_subdomains.txt", "a") # load domain/subdomain list from file URLS = open("subdomains.txt", "r").read().split("\n") URLS_length = len(URLS) # Count the number of live subdomains live_count = 0 # Start the load operations and mark each future with its URL future_to_url = { executor.submit(load_url_http, protocol, url, timeout): url for url in URLS } for i, future in zip(range(URLS_length), concurrent.futures.as_completed(future_to_url)): url = future_to_url[future] print(f"\r--> Checking live subdomain.........{i+1}/{URLS_length}", end="") try: data = future.result() # If `load_url_http` returns any status code if data != None: # print(f'{protocol}://{url}:{data}') live_count = live_count + 1 file_live_subdomain.write(f"\n{protocol}://" + url) except Exception as exc: print(exc) print(f"\n[+] Live domain: {live_count}/{URLS_length}", end="") file_live_subdomain.close() print("\n--- %s seconds ---" % (time.time() - start_time))
运行测试结果
┌──(quangtb㉿QuangTB)-[/mnt/e/DATA/Downloads] └─$ python3 python-request-multil.py Number of worker: 100 --> Checking live subdomain.........1117/1117 [+] Live domain: 344/1117 --- 67.41670227050781 seconds --- ┌──(quangtb㉿QuangTB)-[/mnt/e/DATA/Downloads] └─$ python3 python-request-multil.py Number of worker: 200 --> Checking live subdomain.........1117/1117 [+] Live domain: 344/1117 --- 54.6825795173645 seconds --- ┌──(quangtb㉿QuangTB)-[/mnt/e/DATA/Downloads] └─$ python3 python-request-multil.py Number of worker: 300 --> Checking live subdomain.........1117/1117 [+] Live domain: 339/1117 --- 54.186068058013916 seconds --- ┌──(quangtb㉿QuangTB)-[/mnt/e/DATA/Downloads] └─$ python3 python-request-multil.py Number of worker: 400 --> Checking live subdomain.........1117/1117 [+] Live domain: 344/1117 --- 54.19181728363037 seconds ---
优化方案
1. 切换为异步IO框架
Python线程受全局解释器锁(GIL)限制,IO密集型场景下,异步IO(如aiohttp)能支撑数千级别的并发,不会出现多线程超过阈值后性能瓶颈的问题。这是你增加线程数但速度没提升的核心原因——线程切换开销抵消了并发收益。
2. 复用HTTP连接池
原代码中requests每次请求新建TCP连接,握手开销大。改用requests.Session(多线程场景)或aiohttp.ClientSession(异步场景)复用连接池,能大幅减少连接建立时间。
3. 缩短超时时间
当前10秒超时过长,大部分存活域名响应时间在1-3秒内,将超时设为2-3秒可快速丢弃无响应域名,提升整体检测效率。
4. 批量处理文件IO
原代码每找到一个存活域名就写一次文件,改为先将结果存入列表,最后一次性写入,减少磁盘IO操作次数。
5. 过滤无效域名
提前过滤subdomains.txt中的空行和无效格式域名,减少不必要的请求。
优化后异步版本示例代码
import time import aiohttp import asyncio async def check_live(session, protocol, domain, timeout): url = f"{protocol}://{domain}" try: async with session.head(url, timeout=timeout) as resp: return domain, resp.status except Exception: return domain, None async def main(): start_time = time.time() concurrency = 1000 # 异步可轻松设置数千并发 protocol = "http" timeout = 3 print(f"并发数: {concurrency}") # 读取并过滤域名列表 with open("subdomains.txt", "r") as f: urls = [line.strip() for line in f if line.strip()] total_count = len(urls) live_count = 0 live_domains = [] # 创建带连接池的异步Session connector = aiohttp.TCPConnector(limit=concurrency) async with aiohttp.ClientSession(connector=connector) as session: # 批量创建检测任务 tasks = [check_live(session, protocol, url, timeout) for url in urls] for i, task in enumerate(asyncio.as_completed(tasks), 1): domain, status = await task print(f"\r--> 正在检测... {i}/{total_count}", end="") if status is not None: live_count += 1 live_domains.append(f"{protocol}://{domain}") # 一次性写入存活域名 with open("live_subdomains.txt", "a") as f: if live_domains: f.write("\n".join(live_domains) + "\n") print(f"\n[+] 存活域名: {live_count}/{total_count}") print(f"--- {time.time() - start_time} 秒 ---") if __name__ == "__main__": asyncio.run(main())
优化效果说明
- 异步IO绕过GIL限制,可同时处理数千请求,并发能力远超多线程
- 连接复用减少TCP握手开销,提升请求响应速度
- 更短的超时快速过滤无响应域名,避免无效等待
- 批量文件IO降低磁盘操作延迟
内容的提问来源于stack exchange,提问作者Quang Tran
相关产品推荐
相关产品推荐

