使用Scrapy爬取hypeauditor翻页返回302无数据问题求助
Scrapy翻页请求返回302无法抓取数据的问题解决
问题背景
使用Scrapy爬取hypeauditor的YouTube游戏频道榜单页面,第一页请求返回200状态码并能正常抓取数据,但翻页请求(如?p=2)始终返回302状态码,无法获取有效内容。
爬虫代码
import scrapy from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor from scrapy.crawler import CrawlerProcess class chnlids(scrapy.Spider): name = "chnlids" allowed_domains = ["hypeauditor.com"] start_urls = ["https://hypeauditor.com/top-youtube-video-games/"] handle_httpstatus_list = [302] def start_requests(self): for url in self.start_urls: yield scrapy.Request(url, meta={'dont_redirect':True}) def parse(self,response): i = 0 for key in response.css("div.contributor__content a.contributor::attr(href)").getall(): yield{ "channel_id" : response.css("div.contributor__content a.contributor::attr(href)").get() } i += 1 if(self.page_number != 10): self.page_number += 1 next_page = f'https://hypeauditor.com/top-youtube-video-games/?p={self.page_number}' yield scrapy.Request(next_page, callback=self.parse,meta={'dont_redirect': True}
运行结果
.... 2023-04-30 10:32:03 [scrapy.core.scraper] DEBUG: Scraped from <200 https://hypeauditor.com/top-youtube-video-games/> {'channel_id': '/youtube/UCwIWAbIeu0xI0ReKWOcw3eg/'} 2023-04-30 10:32:04 [scrapy.core.engine] DEBUG: Crawled (302) <GET https://hypeauditor.com/top-youtube-video-games/?p=2> (referer: https://hypeauditor.com/top-youtube-video-games/) /usr/local/lib/python3.10/dist-packages/scrapy/selector/unified.py:83: UserWarning: Selector got both text and root, root is being ignored. .....
已尝试的方法
- 设置
handle_httpstatus_list = [302] - 在Request.meta中添加
meta={'dont_redirect': True} - 在settings.py中添加
REDIRECT_ENABLED = False
解决方案
1. 修复未初始化变量错误
代码中self.page_number未初始化就直接调用,会导致运行报错,需在类中添加初始化逻辑:
def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) self.page_number = 1
2. 处理网站反爬机制
302重定向大概率是网站反爬检测导致,需模拟真实浏览器请求:
- 添加完整请求头:在所有请求中加入
User-Agent等必要头信息,模拟浏览器访问:
self.headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/112.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Referer': 'https://hypeauditor.com/top-youtube-video-games/' }
- 启用Cookies并允许重定向:去掉
dont_redirect参数,让Scrapy自动处理重定向并维护会话Cookies,同时确保settings.py中COOKIES_ENABLED = True。
3. 修正数据提取逻辑
当前循环中每次都提取第一个channel_id,应使用循环变量直接取值:
for channel_href in response.css("div.contributor__content a.contributor::attr(href)").getall(): yield {"channel_id": channel_href}
完整修正后的代码
import scrapy from scrapy.crawler import CrawlerProcess class chnlids(scrapy.Spider): name = "chnlids" allowed_domains = ["hypeauditor.com"] start_urls = ["https://hypeauditor.com/top-youtube-video-games/"] def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) self.page_number = 1 self.headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/112.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Referer': 'https://hypeauditor.com/top-youtube-video-games/' } def start_requests(self): for url in self.start_urls: yield scrapy.Request(url, headers=self.headers) def parse(self, response): # 正确提取每个频道ID for channel_href in response.css("div.contributor__content a.contributor::attr(href)").getall(): yield {"channel_id": channel_href} # 翻页逻辑 if self.page_number < 10: self.page_number += 1 next_page = f'https://hypeauditor.com/top-youtube-video-games/?p={self.page_number}' yield scrapy.Request(next_page, callback=self.parse, headers=self.headers)
内容的提问来源于stack exchange,提问作者Pramess
相关产品推荐
相关产品推荐

