You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Scrapy+Playwright实现无限滚动网站全量爬取?

解决Scrapy+Playwright爬取无限滚动页面的问题

要爬取https://www.futuretools.io这类无限滚动页面,核心是模拟浏览器滚动到底部并等待新内容加载,重复该过程直到没有新内容出现。以下是修改后的完整代码:

from scrapy import Spider
from scrapy_playwright.page import PageMethod


class ToolsSpider(Spider):
    name = "tools"

    def start_requests(self):
        yield scrapy.Request(
            "https://www.futuretools.io",
            meta=dict(
                playwright=True,
                playwright_page_methods=[
                    PageMethod("wait_for_selector", "div.jetboost-list-wrapper-n5zn > div.w-dyn-items div.tool"),
                ],
                # 保留页面实例用于后续操作
                playwright_include_page=True,
            ),
        )

    async def parse(self, response):
        page = response.meta["playwright_page"]
        item_count = 0

        while True:
            # 获取当前已加载的工具元素总数
            current_count = await page.locator("div.jetboost-list-wrapper-n5zn > div.w-dyn-items div.tool").count()
            
            # 元素数量不再增加时,说明已加载全部内容,终止循环
            if current_count == item_count:
                break
            
            item_count = current_count
            
            # 滚动到页面底部
            await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
            # 等待新元素加载,超时时间设为5秒
            await page.wait_for_selector("div.jetboost-list-wrapper-n5zn > div.w-dyn-items div.tool", state="attached", timeout=5000)
            # 给页面预留加载缓冲时间
            await page.wait_for_timeout(1000)

        # 获取完整页面HTML后解析所有元素
        final_html = await page.content()
        selector = scrapy.Selector(text=final_html)
        
        for tool in selector.css("div.jetboost-list-wrapper-n5zn > div.w-dyn-items div.tool"):
            yield {
                'title': tool.css('div.div-block-18 a.tool-item-link---new::text').get(),
                'description': tool.css('div.tool-item-description-box---new::text').get(),
                'total_votes': tool.css('div.list-upvote div.text-block-52::text').get(),
                'category': tool.css('div.collection-list-wrapper-9 div.text-block-53::text').get()
            }
        
        # 关闭页面释放资源
        await page.close()

关键修改说明

  • playwright_include_page=True:在请求元数据中添加该参数,让Scrapy保留Playwright的页面实例,用于后续的滚动、等待操作。
  • 循环滚动逻辑:
    • 每次滚动前记录当前已加载的元素数量,作为判断是否加载完成的依据
    • 通过window.scrollTo模拟浏览器滚动到底部
    • 等待新的工具元素加载完成,确保内容加载完毕
    • 当元素数量不再变化时,停止循环,确认所有内容已加载
  • 统一解析:等全部内容加载完成后,一次性获取完整页面HTML进行解析,避免重复解析部分内容

注意事项

  • 可根据页面实际加载速度,调整wait_for_timeout(毫秒)和wait_for_selector的超时时间
  • 避免滚动过于频繁,防止触发网站反爬机制

内容的提问来源于stack exchange,提问作者Anish Thapa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 01:57:38