You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多线程优化问询:如何高效批量检测存活域名/子域名?

问题描述

我编写了一个程序,用于检测subdomains.txt文件中的域名/子域名是否可通过HTTP/HTTPS协议存活。实现逻辑是向目标发送HEAD请求,若能获取响应状态码则判定为存活(对应load_url_http函数)。

为提升运行速度,我使用了concurrent.futures.ThreadPoolExecutor,初始设置线程数为200,但将线程数提升至300、400后,程序速度并未得到明显提升。我希望对程序进行优化,使其能够同时发送数千次请求。以下是我的程序源码及运行测试结果:

程序源码(python-request-multil.py)

import time

import requests
import concurrent.futures


def load_url_http(protocol: str, domain: str, timeout: int = 10):
        try:
            conn = requests.head(protocol + "://" + domain, timeout=timeout)
            return conn.status_code
        except Exception:
            return None


#--- main ---#
start_time = time.time()

worker = 400
protocol = "http"
timeout = 10

print("Number of worker:", worker)

with concurrent.futures.ThreadPoolExecutor(max_workers=worker) as executor:
    # The file object that the subdomain lives on will be written to
    file_live_subdomain = open("live_subdomains.txt", "a")
    
    # load domain/subdomain list from file
    URLS = open("subdomains.txt", "r").read().split("\n")
    URLS_length = len(URLS)
    
    # Count the number of live subdomains
    live_count = 0
    
    # Start the load operations and mark each future with its URL
    future_to_url = {
        executor.submit(load_url_http, protocol, url, timeout): url for url in URLS
    }
    
    for i, future in zip(range(URLS_length), concurrent.futures.as_completed(future_to_url)):
        url = future_to_url[future]
        print(f"\r-->  Checking live subdomain.........{i+1}/{URLS_length}", end="")
        try:
            data = future.result()
            
            # If `load_url_http` returns any status code
            if data != None:
                # print(f'{protocol}://{url}:{data}')
                live_count = live_count + 1
                file_live_subdomain.write(f"\n{protocol}://" + url)
        except Exception as exc:
            print(exc)
    print(f"\n[+] Live domain: {live_count}/{URLS_length}", end="")
    file_live_subdomain.close()

print("\n--- %s seconds ---" % (time.time() - start_time))

运行测试结果

┌──(quangtb㉿QuangTB)-[/mnt/e/DATA/Downloads]
└─$ python3 python-request-multil.py
Number of worker: 100
-->  Checking live subdomain.........1117/1117
[+] Live domain: 344/1117
--- 67.41670227050781 seconds ---

┌──(quangtb㉿QuangTB)-[/mnt/e/DATA/Downloads]
└─$ python3 python-request-multil.py
Number of worker: 200
-->  Checking live subdomain.........1117/1117
[+] Live domain: 344/1117
--- 54.6825795173645 seconds ---

┌──(quangtb㉿QuangTB)-[/mnt/e/DATA/Downloads]
└─$ python3 python-request-multil.py
Number of worker: 300
-->  Checking live subdomain.........1117/1117
[+] Live domain: 339/1117
--- 54.186068058013916 seconds ---

┌──(quangtb㉿QuangTB)-[/mnt/e/DATA/Downloads]
└─$ python3 python-request-multil.py
Number of worker: 400
-->  Checking live subdomain.........1117/1117
[+] Live domain: 344/1117
--- 54.19181728363037 seconds ---

优化方案

1. 切换为异步IO框架

Python线程受全局解释器锁(GIL)限制,IO密集型场景下,异步IO(如aiohttp)能支撑数千级别的并发,不会出现多线程超过阈值后性能瓶颈的问题。这是你增加线程数但速度没提升的核心原因——线程切换开销抵消了并发收益。

2. 复用HTTP连接池

原代码中requests每次请求新建TCP连接,握手开销大。改用requests.Session(多线程场景)或aiohttp.ClientSession(异步场景)复用连接池,能大幅减少连接建立时间。

3. 缩短超时时间

当前10秒超时过长,大部分存活域名响应时间在1-3秒内,将超时设为2-3秒可快速丢弃无响应域名,提升整体检测效率。

4. 批量处理文件IO

原代码每找到一个存活域名就写一次文件,改为先将结果存入列表,最后一次性写入,减少磁盘IO操作次数。

5. 过滤无效域名

提前过滤subdomains.txt中的空行和无效格式域名,减少不必要的请求。

优化后异步版本示例代码

import time
import aiohttp
import asyncio

async def check_live(session, protocol, domain, timeout):
    url = f"{protocol}://{domain}"
    try:
        async with session.head(url, timeout=timeout) as resp:
            return domain, resp.status
    except Exception:
        return domain, None

async def main():
    start_time = time.time()
    concurrency = 1000  # 异步可轻松设置数千并发
    protocol = "http"
    timeout = 3
    print(f"并发数: {concurrency}")

    # 读取并过滤域名列表
    with open("subdomains.txt", "r") as f:
        urls = [line.strip() for line in f if line.strip()]
    total_count = len(urls)
    live_count = 0
    live_domains = []

    # 创建带连接池的异步Session
    connector = aiohttp.TCPConnector(limit=concurrency)
    async with aiohttp.ClientSession(connector=connector) as session:
        # 批量创建检测任务
        tasks = [check_live(session, protocol, url, timeout) for url in urls]
        for i, task in enumerate(asyncio.as_completed(tasks), 1):
            domain, status = await task
            print(f"\r-->  正在检测... {i}/{total_count}", end="")
            if status is not None:
                live_count += 1
                live_domains.append(f"{protocol}://{domain}")
    
    # 一次性写入存活域名
    with open("live_subdomains.txt", "a") as f:
        if live_domains:
            f.write("\n".join(live_domains) + "\n")
    
    print(f"\n[+] 存活域名: {live_count}/{total_count}")
    print(f"--- {time.time() - start_time} 秒 ---")

if __name__ == "__main__":
    asyncio.run(main())

优化效果说明

  • 异步IO绕过GIL限制,可同时处理数千请求,并发能力远超多线程
  • 连接复用减少TCP握手开销,提升请求响应速度
  • 更短的超时快速过滤无响应域名,避免无效等待
  • 批量文件IO降低磁盘操作延迟

内容的提问来源于stack exchange,提问作者Quang Tran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 05:01:47