You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取hypeauditor翻页返回302无数据问题求助

Scrapy翻页请求返回302无法抓取数据的问题解决

问题背景

使用Scrapy爬取hypeauditor的YouTube游戏频道榜单页面,第一页请求返回200状态码并能正常抓取数据,但翻页请求(如?p=2)始终返回302状态码,无法获取有效内容。

爬虫代码

import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from scrapy.crawler import CrawlerProcess

class chnlids(scrapy.Spider):
    name = "chnlids"
    
    allowed_domains = ["hypeauditor.com"]
    start_urls = ["https://hypeauditor.com/top-youtube-video-games/"]
    handle_httpstatus_list = [302]
    
    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(url, meta={'dont_redirect':True})
    

    def parse(self,response):
      i = 0
      for key in response.css("div.contributor__content a.contributor::attr(href)").getall():

        yield{
            "channel_id" : response.css("div.contributor__content a.contributor::attr(href)").get()
        }
        i += 1
        
      if(self.page_number != 10):
        self.page_number += 1
        next_page = f'https://hypeauditor.com/top-youtube-video-games/?p={self.page_number}'
        yield scrapy.Request(next_page, callback=self.parse,meta={'dont_redirect': True}

运行结果

....
2023-04-30 10:32:03 [scrapy.core.scraper] DEBUG: Scraped from <200 https://hypeauditor.com/top-youtube-video-games/>
{'channel_id': '/youtube/UCwIWAbIeu0xI0ReKWOcw3eg/'}
2023-04-30 10:32:04 [scrapy.core.engine] DEBUG: Crawled (302) <GET https://hypeauditor.com/top-youtube-video-games/?p=2> (referer: https://hypeauditor.com/top-youtube-video-games/)
/usr/local/lib/python3.10/dist-packages/scrapy/selector/unified.py:83: UserWarning: Selector got both text and root, root is being ignored.
.....

已尝试的方法

  • 设置handle_httpstatus_list = [302]
  • 在Request.meta中添加meta={'dont_redirect': True}
  • 在settings.py中添加REDIRECT_ENABLED = False

解决方案

1. 修复未初始化变量错误

代码中self.page_number未初始化就直接调用,会导致运行报错,需在类中添加初始化逻辑:

def __init__(self, *args, **kwargs):
    super().__init__(*args, **kwargs)
    self.page_number = 1

2. 处理网站反爬机制

302重定向大概率是网站反爬检测导致,需模拟真实浏览器请求:

  • 添加完整请求头:在所有请求中加入User-Agent等必要头信息,模拟浏览器访问:
self.headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/112.0.0.0 Safari/537.36',
    'Accept-Language': 'en-US,en;q=0.9',
    'Referer': 'https://hypeauditor.com/top-youtube-video-games/'
}
  • 启用Cookies并允许重定向:去掉dont_redirect参数,让Scrapy自动处理重定向并维护会话Cookies,同时确保settings.py中COOKIES_ENABLED = True。

3. 修正数据提取逻辑

当前循环中每次都提取第一个channel_id,应使用循环变量直接取值:

for channel_href in response.css("div.contributor__content a.contributor::attr(href)").getall():
    yield {"channel_id": channel_href}

完整修正后的代码

import scrapy
from scrapy.crawler import CrawlerProcess

class chnlids(scrapy.Spider):
    name = "chnlids"
    
    allowed_domains = ["hypeauditor.com"]
    start_urls = ["https://hypeauditor.com/top-youtube-video-games/"]
    
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.page_number = 1
        self.headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/112.0.0.0 Safari/537.36',
            'Accept-Language': 'en-US,en;q=0.9',
            'Referer': 'https://hypeauditor.com/top-youtube-video-games/'
        }
    
    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(url, headers=self.headers)
    
    def parse(self, response):
        # 正确提取每个频道ID
        for channel_href in response.css("div.contributor__content a.contributor::attr(href)").getall():
            yield {"channel_id": channel_href}
        
        # 翻页逻辑
        if self.page_number < 10:
            self.page_number += 1
            next_page = f'https://hypeauditor.com/top-youtube-video-games/?p={self.page_number}'
            yield scrapy.Request(next_page, callback=self.parse, headers=self.headers)

内容的提问来源于stack exchange,提问作者Pramess

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 03:05:37