Scrapy Playwright滚动加载问题:仅获48条数据而非全部630+
问题
我有一个页面,初始仅显示48个产品,向下滚动会自动加载更多内容,总计约有630个产品。但我的Scrapy爬虫始终只能获取48条结果,无法获取全部内容,请问原因是什么?应该修改哪些部分?
页面链接:https://www.tradeinn.com/runnerinn/en/mens-shoes-trail-running-shoes/10005/s#fq=id_familia=10002&sort=v30_sum;desc@tm10;asc&fe=&pf=id_subfamilia=10005@&start=0
我的爬虫代码如下:
import scrapy from scrapy_playwright.page import PageMethod class PicturesSpider(scrapy.Spider): name = 'pictures' allowed_domains = ['www.tradeinn.com'] start_urls = ['http://www.tradeinn.com/'] def start_requests(self): yield scrapy.Request(url='https://www.tradeinn.com/runnerinn/en/mens-shoes-trail-running-shoes/10005/s#fq=id_familia=10002&sort=v30_sum;desc@tm10;asc&fe=&pf=id_subfamilia=10005&&start=144', meta={'playwright': True, 'playwright_include_page': True, 'playwright_page_method': [PageMethod('wait_for_selector', 'div::boton_cargar_mas.color_runnerinn'), PageMethod("evaluate", "window.scrollBy(0, document.body.scrollHeight)")]}, callback=self.parse) def parse(self, response): images = response.css('div.BoxImage') for image in images: image_link = image.css('img::attr(src)').get() image_description = image.css('img::attr(alt)').get() yield { 'image_link': image_link, 'image_description': image_description }
原因分析
- 单次滚动无法触发全部加载:当前代码只执行了一次滚动操作,而页面需要多次滚动才能加载完所有630+产品,单次操作最多只能加载第一批后的少量内容,覆盖不了全部。
- 等待选择器语法错误:
div::boton_cargar_mas.color_runnerinn是错误的选择器写法,::是伪元素选择器,这里应该用类选择器.,正确写法应为div.boton_cargar_mas.color_runnerinn(如果按钮确实是这个类)。错误的选择器会导致等待逻辑失效,页面还没加载更多内容就开始解析。 - 起始URL参数错误:你用的URL里
start=144,直接跳过了前面的产品,初始加载的就是对应start值的48条,不是从第一个产品开始爬取。
修改方案
1. 核心修改点
- 实现循环滚动+等待,直到页面不再加载新内容;
- 修正选择器语法,确保等待逻辑生效;
- 使用正确的初始URL(
start=0)从第一个产品开始爬取。
完整修改后的代码
import scrapy from scrapy_playwright.page import PageMethod from playwright.async_api import Page from scrapy.selector import Selector class PicturesSpider(scrapy.Spider): name = 'pictures' allowed_domains = ['www.tradeinn.com'] def start_requests(self): # 使用初始start=0的URL,从第一个产品开始爬取 url = 'https://www.tradeinn.com/runnerinn/en/mens-shoes-trail-running-shoes/10005/s#fq=id_familia=10002&sort=v30_sum;desc@tm10;asc&fe=&pf=id_subfamilia=10005@&start=0' yield scrapy.Request( url=url, meta={ 'playwright': True, 'playwright_include_page': True, 'playwright_page_methods': [ PageMethod('wait_for_selector', 'div.BoxImage') # 等待初始产品列表加载完成 ] }, callback=self.parse ) async def parse(self, response): page: Page = response.meta['playwright_page'] last_page_height = await page.evaluate("document.body.scrollHeight") while True: # 滚动到当前页面底部 await page.evaluate("window.scrollTo(0, document.body.scrollHeight)") # 等待新内容加载(可根据页面实际加载速度调整等待时间,或替换为更精准的元素等待) await page.wait_for_timeout(2000) # 获取滚动后的页面高度 new_page_height = await page.evaluate("document.body.scrollHeight") # 如果页面高度不再变化,说明没有更多内容可加载,退出循环 if new_page_height == last_page_height: break last_page_height = new_page_height # 加载完成后,获取完整页面HTML并解析 full_page_content = await page.content() sel = Selector(text=full_page_content) images = sel.css('div.BoxImage') for image in images: image_link = image.css('img::attr(src)').get() image_description = image.css('img::attr(alt)').get() yield { 'image_link': image_link, 'image_description': image_description } # 关闭Playwright页面 await page.close()
可选优化:用加载按钮代替滚动
如果页面有明确的“加载更多”按钮,可将循环滚动改为循环点击按钮,逻辑更精准:
# 替换parse方法中的循环部分 while True: try: # 定位加载更多按钮,替换为实际的选择器 load_more_btn = page.locator('div.boton_cargar_mas.color_runnerinn') await load_more_btn.click() await page.wait_for_timeout(1500) # 等待内容加载 except: # 按钮不存在,说明加载完成,退出循环 break
内容的提问来源于stack exchange,提问作者PetrSevcik
相关产品推荐
相关产品推荐

