网站检测到异常流量时,如何继续爬虫数据采集?
问题描述
我出于兴趣开发了一个网页爬虫,目标是爬取中文网站9game.cn的游戏缩略图,实现逻辑为:
- 请求目标URL
- 递增URL切换页面
- 查找DOM中的缩略图元素
- 将缩略图保存至本地桌面
我设置了嵌套的if-else语句处理页面不存在或无缩略图的异常情况,此时会打印对应页码并继续执行。但运行一段时间后,即使页面存在缩略图,爬虫也无法保存图片,且收到异常流量检测的警告,希望得到解决该问题的技术帮助。
原代码
from pickle import NONE import requests import urllib.request from bs4 import BeautifulSoup localfile = "C:/Users/MYNAME/Desktop/Chinese Games/" url = "https://www.9game.cn/xiazai/" for x in range(1, 500): page = requests.get(url + str(x) + "/") # Request url and iterate with x soup = BeautifulSoup(page.content, 'lxml') image = soup.find(class_="d-headgame-icon") # Finds the HTML elements that holds the image if image != None: result = image.find("img").attrs['src'] # Extracts the URL of the image from all other elements if result !="": title = image.find("img").attrs['alt'] # Extracts the name of the image from all other elements urllib.request.urlretrieve(result, localfile + title + ".jpg") # Saves images to Desktop as a JPG file else: print(x) # Prints the page number if there is no image else: print(x) # Prints the page number if there is no image
解决方案
触发异常流量检测的核心原因是爬虫行为与普通用户差异过大,以下是针对性优化措施:
添加浏览器请求头
网站会通过请求头识别非浏览器请求,模拟主流浏览器的请求头可降低拦截概率:headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'https://www.9game.cn/' } # 在请求时传入headers page = requests.get(url + str(x) + "/", headers=headers)添加随机请求延迟
无间隔连续请求是典型的爬虫特征,添加随机延迟模拟用户浏览间隔:import random import time # 每次请求后添加1-3秒的随机延迟 time.sleep(random.uniform(1, 3))使用会话保持连接
用requests.Session()维持会话,模拟用户连续访问的状态,避免每次请求都建立新连接:session = requests.Session() session.headers.update(headers) # 统一设置会话的请求头 # 循环中用session.get替代requests.get page = session.get(url + str(x) + "/")替换图片下载方式
urllib.request.urlretrieve的行为容易被识别,改用requests配合会话下载,同时处理标题中的特殊字符避免保存失败:import os from urllib.parse import urljoin # 处理文件名中的非法字符 def sanitize_filename(filename): return "".join([c for c in filename if c not in r'\/:*?"<>|']) # 替换原下载逻辑 img_url = urljoin(url, result) # 处理相对路径的图片链接 img_response = session.get(img_url) if img_response.status_code == 200: sanitized_title = sanitize_filename(title) with open(os.path.join(localfile, f"{sanitized_title}.jpg"), 'wb') as f: f.write(img_response.content)增强异常处理
原代码缺少请求失败的容错逻辑,添加状态码判断和异常捕获,避免无效请求或程序崩溃:try: page = session.get(url + str(x) + "/", timeout=10) page.raise_for_status() # 触发HTTP错误异常 soup = BeautifulSoup(page.content, 'lxml') # 后续解析逻辑... except requests.exceptions.RequestException as e: print(f"页面{x}请求失败: {str(e)}") continue
内容的提问来源于stack exchange,提问作者PolarShock
相关产品推荐
相关产品推荐

