You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Playwright滚动加载问题:仅获48条数据而非全部630+

问题

我有一个页面,初始仅显示48个产品,向下滚动会自动加载更多内容,总计约有630个产品。但我的Scrapy爬虫始终只能获取48条结果,无法获取全部内容,请问原因是什么?应该修改哪些部分?

页面链接:https://www.tradeinn.com/runnerinn/en/mens-shoes-trail-running-shoes/10005/s#fq=id_familia=10002&sort=v30_sum;desc@tm10;asc&fe=&pf=id_subfamilia=10005@&start=0

我的爬虫代码如下:

import scrapy
from scrapy_playwright.page import PageMethod


class PicturesSpider(scrapy.Spider):
    name = 'pictures'
    allowed_domains = ['www.tradeinn.com']
    start_urls = ['http://www.tradeinn.com/']

    def start_requests(self):
        yield scrapy.Request(url='https://www.tradeinn.com/runnerinn/en/mens-shoes-trail-running-shoes/10005/s#fq=id_familia=10002&sort=v30_sum;desc@tm10;asc&fe=&pf=id_subfamilia=10005&&start=144',
                             meta={'playwright': True,
                                   'playwright_include_page': True,
                                   'playwright_page_method': [PageMethod('wait_for_selector', 'div::boton_cargar_mas.color_runnerinn'),
                                                              PageMethod("evaluate", "window.scrollBy(0, document.body.scrollHeight)")]},
                             callback=self.parse)


    def parse(self, response):
        images = response.css('div.BoxImage')
        for image in images:
            image_link = image.css('img::attr(src)').get()
            image_description = image.css('img::attr(alt)').get()
            yield {
                'image_link': image_link,
                'image_description': image_description
            }
原因分析
  • 单次滚动无法触发全部加载:当前代码只执行了一次滚动操作,而页面需要多次滚动才能加载完所有630+产品,单次操作最多只能加载第一批后的少量内容,覆盖不了全部。
  • 等待选择器语法错误:div::boton_cargar_mas.color_runnerinn是错误的选择器写法,::是伪元素选择器,这里应该用类选择器.,正确写法应为div.boton_cargar_mas.color_runnerinn(如果按钮确实是这个类)。错误的选择器会导致等待逻辑失效,页面还没加载更多内容就开始解析。
  • 起始URL参数错误:你用的URL里start=144,直接跳过了前面的产品,初始加载的就是对应start值的48条,不是从第一个产品开始爬取。
修改方案

1. 核心修改点

  • 实现循环滚动+等待,直到页面不再加载新内容;
  • 修正选择器语法,确保等待逻辑生效;
  • 使用正确的初始URL(start=0)从第一个产品开始爬取。

完整修改后的代码

import scrapy
from scrapy_playwright.page import PageMethod
from playwright.async_api import Page
from scrapy.selector import Selector


class PicturesSpider(scrapy.Spider):
    name = 'pictures'
    allowed_domains = ['www.tradeinn.com']

    def start_requests(self):
        # 使用初始start=0的URL,从第一个产品开始爬取
        url = 'https://www.tradeinn.com/runnerinn/en/mens-shoes-trail-running-shoes/10005/s#fq=id_familia=10002&sort=v30_sum;desc@tm10;asc&fe=&pf=id_subfamilia=10005@&start=0'
        yield scrapy.Request(
            url=url,
            meta={
                'playwright': True,
                'playwright_include_page': True,
                'playwright_page_methods': [
                    PageMethod('wait_for_selector', 'div.BoxImage')  # 等待初始产品列表加载完成
                ]
            },
            callback=self.parse
        )

    async def parse(self, response):
        page: Page = response.meta['playwright_page']
        last_page_height = await page.evaluate("document.body.scrollHeight")

        while True:
            # 滚动到当前页面底部
            await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
            # 等待新内容加载(可根据页面实际加载速度调整等待时间,或替换为更精准的元素等待)
            await page.wait_for_timeout(2000)
            # 获取滚动后的页面高度
            new_page_height = await page.evaluate("document.body.scrollHeight")
            
            # 如果页面高度不再变化,说明没有更多内容可加载,退出循环
            if new_page_height == last_page_height:
                break
            last_page_height = new_page_height

        # 加载完成后,获取完整页面HTML并解析
        full_page_content = await page.content()
        sel = Selector(text=full_page_content)

        images = sel.css('div.BoxImage')
        for image in images:
            image_link = image.css('img::attr(src)').get()
            image_description = image.css('img::attr(alt)').get()
            yield {
                'image_link': image_link,
                'image_description': image_description
            }

        # 关闭Playwright页面
        await page.close()

可选优化:用加载按钮代替滚动

如果页面有明确的“加载更多”按钮,可将循环滚动改为循环点击按钮,逻辑更精准:

# 替换parse方法中的循环部分
while True:
    try:
        # 定位加载更多按钮,替换为实际的选择器
        load_more_btn = page.locator('div.boton_cargar_mas.color_runnerinn')
        await load_more_btn.click()
        await page.wait_for_timeout(1500)  # 等待内容加载
    except:
        # 按钮不存在,说明加载完成,退出循环
        break

内容的提问来源于stack exchange,提问作者PetrSevcik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 02:05:32