You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

异步网页爬取中IP轮询实现及防IP封禁问题咨询

Hey there! 作为编程新手遇到IP封禁太正常啦,我来一步步教你把随机代理融入异步代码,还会给你几个简单有效的防封禁技巧,保证不破坏异步的效率~

一、实现随机代理的核心修改

首先你要明确:aiohttp给请求加代理超级简单,不用搞复杂的TCPConnector(那是绑定本地IP用的,你需要的是用代理IP发请求)。只需要在session.get()里加proxy参数就行,每次请求随机选一个代理列表里的IP就OK。

步骤1:准备代理列表

假设你已经有了可用的代理,格式要正确,比如:

import random
import aiohttp
import asyncio

# 假设这是你已经生成好的代理列表,格式为 "http://ip:端口" 或者 "https://ip:端口"
PROXIES = [
    "http://123.45.67.89:8080",
    "http://98.76.54.32:3128",
    # 更多代理...
]

步骤2:修改download_file函数

每次请求前随机选一个代理,还要加个简单的异常处理(避免某个代理失效导致整个任务挂掉):

async def download_file(url):
    # 随机选一个代理
    proxy = random.choice(PROXIES)
    # 模拟浏览器请求头,避免被识别为爬虫
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    try:
        async with aiohttp.ClientSession() as session:
            # 把proxy参数传入请求
            async with session.get(url, proxy=proxy, headers=headers) as resp:
                # 先检查响应状态码,确保请求成功
                if resp.status == 200:
                    content = await resp.read()
                    return content
                else:
                    print(f"请求失败,状态码:{resp.status},URL:{url},代理:{proxy}")
                    return None
    except Exception as e:
        print(f"请求出错:{str(e)},URL:{url},代理:{proxy}")
        # 可以在这里加重试逻辑,比如换个代理再试一次
        return None

步骤3:调整任务逻辑,处理空内容

因为代理可能失效,所以write_file要判断内容是否为空,避免生成空文件:

async def write_file(n, content):
    if content is None:
        print(f"跳过空文件:sync_{n}.html")
        return
    filename = f'sync_{n}.html'
    with open(filename, 'wb') as f:
        f.write(content)
二、保证异步优势的关键细节

你担心同步的随机选代理会破坏异步?完全不用怕!random.choice()只是从列表里挑个字符串,耗时微乎其微,根本不会阻塞异步循环。真正影响异步效率的是阻塞IO操作(比如同步的数据库读写、长时间的计算),这种小操作完全没问题。

另外,如果你想进一步优化,可以用asyncio.Semaphore限制并发请求数,避免一下子发几百个请求把代理或者目标网站打崩,这也是防封禁的关键:

async def main():
    # 限制最多同时10个请求,可以根据情况调整
    semaphore = asyncio.Semaphore(10)
    
    async def bounded_scrape_task(n, url):
        async with semaphore:
            content = await download_file(url)
            await write_file(n, content)
    
    tasks = []
    for n, url in enumerate(open('links.txt').readlines()):
        # 去掉URL里的换行符
        url = url.strip()
        if url:  # 跳过空行
            tasks.append(bounded_scrape_task(n, url))
    await asyncio.gather(*tasks)  # 用gather比wait更直观,新手友好
三、新手友好的额外防封禁技巧

除了代理,这几个简单技巧能大大降低被封的概率:

  • 随机请求头:不要一直用同一个User-Agent,可以搞个列表随机选
  • 随机延迟:每个请求前加一点随机等待,比如await asyncio.sleep(random.uniform(0.3, 1.5)),不要太长,不然影响效率
  • 重试机制:如果某个请求失败,换个代理再试1-2次,比如在download_file里加个循环重试
  • 避免请求高峰:如果目标网站有访问高峰,尽量避开(比如凌晨爬)
完整的优化后代码

把上面的内容整合起来,就是完整的可运行代码:

import random
import aiohttp
import asyncio

# 代理列表(替换成你自己的)
PROXIES = [
    "http://123.45.67.89:8080",
    "http://98.76.54.32:3128",
    # 更多代理...
]

# 可选:多个User-Agent随机选
USER_AGENTS = [
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
    "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.5 Safari/605.1.15",
    "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
]

async def download_file(url):
    for _ in range(2):  # 重试2次
        proxy = random.choice(PROXIES)
        headers = {
            "User-Agent": random.choice(USER_AGENTS)
        }
        try:
            async with aiohttp.ClientSession() as session:
                async with session.get(url, proxy=proxy, headers=headers) as resp:
                    if resp.status == 200:
                        # 加一点随机延迟,避免请求太密集
                        await asyncio.sleep(random.uniform(0.3, 1.0))
                        content = await resp.read()
                        return content
                    else:
                        print(f"状态码异常:{resp.status},URL:{url},代理:{proxy}")
        except Exception as e:
            print(f"请求失败:{str(e)},URL:{url},代理:{proxy}")
        # 重试前换个代理,加个小延迟
        await asyncio.sleep(0.5)
    print(f"两次重试都失败,放弃URL:{url}")
    return None

async def write_file(n, content):
    if content is None:
        print(f"跳过空文件:sync_{n}.html")
        return
    filename = f'sync_{n}.html'
    with open(filename, 'wb') as f:
        f.write(content)

async def main():
    semaphore = asyncio.Semaphore(10)  # 限制并发数
    
    async def bounded_scrape_task(n, url):
        async with semaphore:
            content = await download_file(url)
            await write_file(n, content)
    
    tasks = []
    with open('links.txt', 'r', encoding='utf-8') as f:
        for n, url in enumerate(f.readlines()):
            url = url.strip()
            if url:
                tasks.append(bounded_scrape_task(n, url))
    
    await asyncio.gather(*tasks)

if __name__ == '__main__':
    asyncio.run(main())

内容的提问来源于stack exchange,提问作者TNoms

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 09:32:32