如何使用scrapy_playwright点击无href元素后爬取弹出页面数据
scrapy-playwright 点击多元素爬取弹窗数据解决方案
问题根因
你当前代码只能爬取第一个链接的核心原因:
- 预定义的
PageCoroutine("click", selector="h3.result-heading")默认只会点击匹配到的第一个元素,没有遍历所有可点击条目 - 没有监听点击后触发的新标签页事件,无法获取弹出页的响应数据
实现代码
爬虫代码调整
import scrapy from scrapy_playwright.page import PageCoroutine class PwspiderSpider(scrapy.Spider): name = 'demoo' def start_requests(self): yield scrapy.Request( url="https://boston.craigslist.org/search/npo", meta=dict( playwright=True, playwright_include_page=True ), ) async def parse(self, response): # 从meta中获取当前列表页的page对象 page = response.meta["playwright_page"] # 等待所有可点击标题加载完成 await page.wait_for_selector("h3.result-heading", timeout=10000) # 获取所有标题元素 title_elements = await page.query_selector_all("h3.result-heading") for elem in title_elements: try: # 先监听新弹窗事件,再点击元素 popup_promise = page.wait_for_event("popup", timeout=10000) await elem.click() # 获取弹出页对象 popup_page = await popup_promise # 等待弹出页DOM加载完成 await popup_page.wait_for_load_state("domcontentloaded", timeout=10000) # 提取弹出页数据 body = await popup_page.locator("section#postingbody").all_text_contents() yield { "body": [t.strip() for t in body if t.strip()] } # 关闭弹出页,释放资源 await popup_page.close() # 可选:添加短延迟,避免触发反爬 await page.wait_for_timeout(1000) except Exception as e: self.logger.error(f"处理条目出错: {str(e)}") continue # 全部处理完成后关闭列表页 await page.close()
配置补充(settings.py)
确保项目settings.py中已添加scrapy-playwright必要配置:
# 启用playwright下载处理器 DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } # 指定异步 reactor TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor" # 可选配置:设置浏览器启动参数,比如无头模式、代理等 PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": False, # 调试阶段可以设为False看浏览器操作过程 "slow_mo": 500 }
注意事项
- 如果点击后是当前页跳转而非新弹窗,不需要监听
popup事件,点击后直接等待当前页加载,提取数据后调用page.go_back()返回列表页继续处理下一条即可 - 可根据目标站点反爬策略调整等待时长、请求间隔,避免被拦截
- 可以添加翻页逻辑,爬取多页列表的所有条目数据
内容的提问来源于stack exchange,提问作者an huynh
相关产品推荐
相关产品推荐

