Scrapy Playwright爬取Lidl商品失败:点击加载更多后返回空JSON
问题分析与解决方案
原代码存在的核心问题
- 加载更多按钮判断逻辑错误:
page.locator()始终返回Locator对象,无法通过if pagination_buttons:判断按钮是否存在,必须通过可见性检查来确认。 - 错误等待页面跳转:点击“加载更多”是异步AJAX加载内容,不会触发页面导航,
await page.wait_for_navigation()会超时阻塞后续流程。 - 未循环加载全部商品:原代码仅遍历一次按钮,无法实现“点击直到按钮消失”的需求,导致只加载了初始页的商品。
- 详情页数据提取缺陷:直接对
None调用str()会得到字符串'None',而非预期的None值,影响数据准确性。
修正后的完整代码
import scrapy import datetime from scrapy.crawler import CrawlerProcess from scrapy_playwright.page import PageMethod from scrapy.selector import Selector class LidlSpider(scrapy.Spider): name = 'lidl_snacks' allowed_domains = ['sortiment.lidl.ch'] custom_settings = { 'ROBOTSTXT_OBEY': False, # 配置Playwright下载处理器 'DOWNLOAD_HANDLERS': { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", }, 'PLAYWRIGHT_LAUNCH_OPTIONS': { 'headless': False, # 调试阶段可设为False查看浏览器行为,上线后改回True 'args': ['--start-maximized'], # 最大化窗口避免元素被遮挡 }, } start_urls = [ 'https://sortiment.lidl.ch/de/sussigkeiten-snacks#/', ] def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, dont_filter=True, callback=self.parse, meta={ 'playwright': True, 'playwright_include_page': True, 'playwright_page_methods':[ PageMethod('wait_for_selector', 'div.product-item-info', state='attached'), ] } ) async def scroll_to_bottom(self, page): await page.evaluate("window.scrollTo(0, document.body.scrollHeight);") async def parse(self, response): page = response.meta["playwright_page"] # 循环点击加载更多直到按钮消失 while True: load_more_btn = page.locator("button.primary.amscroll-load-button-new") try: # 等待按钮可见,超时5秒则视为无更多内容 await load_more_btn.wait_for(state="visible", timeout=5000) # 点击按钮前先滚动到按钮位置,避免被遮挡 await load_more_btn.scroll_into_view_if_needed() await load_more_btn.click() # 等待新商品加载完成,超时10秒 await page.wait_for_selector("div.product-item-info", state="attached", timeout=10000) # 滚动到底部确保所有内容加载 await self.scroll_to_bottom(page) # 等待1秒让AJAX请求完成 await page.wait_for_timeout(1000) except: # 按钮不可见或超时,退出循环 self.logger.info("所有商品已加载完成") break # 提取页面所有商品链接 content = await page.content() sel = Selector(text=content) produkte = sel.css('div.product-item-info') self.logger.info(f"共抓取到 {len(produkte)} 个商品") for produkt in produkte: produkt_url = produkt.css('a.product-item-link::attr(href)').get() if produkt_url: # 补全相对路径为完整URL if not produkt_url.startswith('http'): produkt_url = f"https://sortiment.lidl.ch{produkt_url}" yield scrapy.Request( produkt_url, callback=self.parse_produkt, meta={'url': response.meta['url']} ) # 关闭页面释放资源 await page.close() def parse_produkt(self, response): # 提取数据时处理空值,避免出现'None'字符串 brand = response.css('p.brand-name::text').get() detail = response.css('span.base::text').get() actual_price = response.css('strong.pricefield__price::attr(content)').get() mini_dict = { 'retailer': self.name, 'datetime': datetime.date.today(), 'categorie': None, 'id': None, 'brand': brand.strip() if brand else None, 'detail': detail.strip() if detail else None, 'actual_price': actual_price, 'quantity': None, 'regular_price': None, 'price_per_unit': None, } yield mini_dict if __name__ == "__main__": process = CrawlerProcess() process.crawl(LidlSpider) process.start()
关键修正说明
- Playwright配置补充:在
custom_settings中添加了必要的下载处理器和浏览器启动参数,确保Playwright正常工作。 - 循环加载逻辑:通过
while True循环+异常捕获,实现“点击加载更多直到按钮消失”的需求,每次点击后等待新商品加载完成。 - 元素可见性与滚动处理:点击按钮前先滚动到按钮位置,避免因元素被遮挡导致点击失败;每次加载后滚动到底部确保内容完整加载。
- 空值处理优化:提取详情页数据时,先判断值是否存在,再做strip处理,避免将
None转为字符串'None'。 - URL补全:处理商品详情页的相对路径,转为完整URL确保请求有效。
内容的提问来源于stack exchange,提问作者user24820468
相关产品推荐
相关产品推荐

