You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网站检测到异常流量时,如何继续爬虫数据采集?

问题描述

我出于兴趣开发了一个网页爬虫,目标是爬取中文网站9game.cn的游戏缩略图,实现逻辑为:

  1. 请求目标URL
  2. 递增URL切换页面
  3. 查找DOM中的缩略图元素
  4. 将缩略图保存至本地桌面

我设置了嵌套的if-else语句处理页面不存在或无缩略图的异常情况,此时会打印对应页码并继续执行。但运行一段时间后,即使页面存在缩略图,爬虫也无法保存图片,且收到异常流量检测的警告,希望得到解决该问题的技术帮助。

原代码

from pickle import NONE
import requests
import urllib.request
from bs4 import BeautifulSoup


localfile = "C:/Users/MYNAME/Desktop/Chinese Games/"
url = "https://www.9game.cn/xiazai/" 


for x in range(1, 500):
        page = requests.get(url + str(x) + "/")         # Request url and iterate with x 
        soup = BeautifulSoup(page.content, 'lxml') 
        image = soup.find(class_="d-headgame-icon")     # Finds the HTML elements that holds the image 
        if image != None:                           
            result = image.find("img").attrs['src']     # Extracts the URL of the image from all other elements
            if result !="":
                title = image.find("img").attrs['alt']  # Extracts the name of the image from all other elements
                urllib.request.urlretrieve(result, localfile + title + ".jpg")  # Saves images to Desktop as a JPG file
            else:
                print(x)    # Prints the page number if there is no image
        else:
            print(x)        # Prints the page number if there is no image
解决方案

触发异常流量检测的核心原因是爬虫行为与普通用户差异过大,以下是针对性优化措施:

  • 添加浏览器请求头
    网站会通过请求头识别非浏览器请求,模拟主流浏览器的请求头可降低拦截概率:

    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Referer': 'https://www.9game.cn/'
    }
    # 在请求时传入headers
    page = requests.get(url + str(x) + "/", headers=headers)
    
  • 添加随机请求延迟
    无间隔连续请求是典型的爬虫特征,添加随机延迟模拟用户浏览间隔:

    import random
    import time
    
    # 每次请求后添加1-3秒的随机延迟
    time.sleep(random.uniform(1, 3))
    
  • 使用会话保持连接
    用requests.Session()维持会话,模拟用户连续访问的状态,避免每次请求都建立新连接:

    session = requests.Session()
    session.headers.update(headers)  # 统一设置会话的请求头
    
    # 循环中用session.get替代requests.get
    page = session.get(url + str(x) + "/")
    
  • 替换图片下载方式
    urllib.request.urlretrieve的行为容易被识别,改用requests配合会话下载,同时处理标题中的特殊字符避免保存失败:

    import os
    from urllib.parse import urljoin
    
    # 处理文件名中的非法字符
    def sanitize_filename(filename):
        return "".join([c for c in filename if c not in r'\/:*?"<>|'])
    
    # 替换原下载逻辑
    img_url = urljoin(url, result)  # 处理相对路径的图片链接
    img_response = session.get(img_url)
    if img_response.status_code == 200:
        sanitized_title = sanitize_filename(title)
        with open(os.path.join(localfile, f"{sanitized_title}.jpg"), 'wb') as f:
            f.write(img_response.content)
    
  • 增强异常处理
    原代码缺少请求失败的容错逻辑,添加状态码判断和异常捕获,避免无效请求或程序崩溃:

    try:
        page = session.get(url + str(x) + "/", timeout=10)
        page.raise_for_status()  # 触发HTTP错误异常
        soup = BeautifulSoup(page.content, 'lxml')
        # 后续解析逻辑...
    except requests.exceptions.RequestException as e:
        print(f"页面{x}请求失败: {str(e)}")
        continue
    

内容的提问来源于stack exchange,提问作者PolarShock

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 23:42:29