Scrapy爬取DiscoverUni动态页面:分页导航问题求助
解决动态加载页面的Scrapy分页爬取问题
针对你遇到的动态加载分页无XHR JSON数据的情况,给你两种可行的解决思路及代码实现:
思路一:排查页面分页请求逻辑(优先尝试)
先通过浏览器开发者工具定位分页触发的真实请求:
- 打开目标页面,开启F12的Network面板
- 点击「下一页」按钮,观察Network中的请求变化:
- 若发现GET请求URL带有
page/offset等参数,直接构造分页URL请求 - 若为POST请求,提取页面中的表单参数(如
__RequestVerificationToken、分页页码等),构造FormRequest
- 若发现GET请求URL带有
代码示例(假设为POST分页请求)
class UnispiderSpider(scrapy.Spider): name = 'unispider' start_urls = ['https://www.discoveruni.gov.uk/course-finder/results/'] base_url = 'https://www.discoveruni.gov.uk/course-finder/results/' current_page = 1 def parse(self, response): # 提取课程数据(保留原有逻辑) course_list = response.xpath( '//div[@class="course-finder-results__result-accordion-body-content comparison-course-area mb-4"]') for course in course_list: courseidentifier = course.xpath('@data-courseidentifier').get() uniname = course.xpath('@data-uniname').get() uniid = course.xpath('@data-uniid').get() coursename = course.xpath('@data-coursename').get() link = course.xpath('a/@href').get() yield { 'courseidentifier': courseidentifier, 'uniname': uniname, 'uniid': uniid, 'coursename': coursename, 'link': link } # 提取分页所需的验证令牌(从页面隐藏元素获取) token = response.xpath('//input[@name="__RequestVerificationToken"]/@value').get() if not token: return # 构造下一页请求 self.current_page += 1 # 可通过页面元素判断是否还有下一页(比如检查下一页按钮是否禁用) has_next = response.xpath('//button[contains(@class, "pagination__next") and not(@disabled)]') if has_next: yield scrapy.FormRequest( url=self.base_url, formdata={ '__RequestVerificationToken': token, 'page': str(self.current_page), # 补充页面中其他必要的表单参数(如搜索条件等,从Network中复制) }, callback=self.parse )
思路二:用Playwright模拟浏览器翻页(纯前端渲染场景)
如果页面完全通过JS渲染、无后台接口返回数据,直接用Scrapy结合Playwright模拟浏览器点击翻页:
步骤1:安装依赖
pip install scrapy-playwright
步骤2:修改Scrapy配置(settings.py)
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, # 测试时可改为False查看浏览器操作 "timeout": 30000, } USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
步骤3:修改Spider代码
from scrapy_playwright.page import PageCoroutine class UnispiderSpider(scrapy.Spider): name = 'unispider' start_urls = ['https://www.discoveruni.gov.uk/course-finder/results/'] def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, meta={ "playwright": True, "playwright_page_coroutines": [ PageCoroutine("wait_for_selector", "div.course-finder-results__result-accordion-body-content"), ], }, callback=self.parse ) def parse(self, response): # 提取课程数据(保留原有逻辑) course_list = response.xpath( '//div[@class="course-finder-results__result-accordion-body-content comparison-course-area mb-4"]') for course in course_list: courseidentifier = course.xpath('@data-courseidentifier').get() uniname = course.xpath('@data-uniname').get() uniid = course.xpath('@data-uniid').get() coursename = course.xpath('@data-coursename').get() link = course.xpath('a/@href').get() yield { 'courseidentifier': courseidentifier, 'uniname': uniname, 'uniid': uniid, 'coursename': coursename, 'link': link } # 模拟点击下一页按钮(需根据页面实际元素调整选择器) next_button_selector = 'button[contains(@class, "pagination__next")]' has_next = response.xpath(next_button_selector + '[not(@disabled)]') if has_next: yield scrapy.Request( response.url, meta={ "playwright": True, "playwright_page_coroutines": [ PageCoroutine("click", next_button_selector), PageCoroutine("wait_for_selector", "div.course-finder-results__result-accordion-body-content"), ], "playwright_continue_page": True, # 复用当前浏览器页面 }, callback=self.parse )
注意事项
- 若使用思路一,务必仔细核对Network中的请求参数,遗漏参数会导致请求失败
- 思路二中的按钮选择器需根据页面实际HTML结构调整,可通过浏览器元素审查工具获取准确选择器
- 避免高频请求,可在settings.py中添加
DOWNLOAD_DELAY = 2降低反爬风险
内容的提问来源于stack exchange,提问作者Jason_D
相关产品推荐
相关产品推荐

