You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用scrapy_playwright点击无href元素后爬取弹出页面数据

scrapy-playwright 点击多元素爬取弹窗数据解决方案

问题根因

你当前代码只能爬取第一个链接的核心原因:

  1. 预定义的PageCoroutine("click", selector="h3.result-heading")默认只会点击匹配到的第一个元素,没有遍历所有可点击条目
  2. 没有监听点击后触发的新标签页事件,无法获取弹出页的响应数据

实现代码

爬虫代码调整

import scrapy
from scrapy_playwright.page import PageCoroutine

class PwspiderSpider(scrapy.Spider):
    name = 'demoo'

    def start_requests(self):
        yield scrapy.Request(
            url="https://boston.craigslist.org/search/npo",
            meta=dict(
                playwright=True,
                playwright_include_page=True
            ),
        )

    async def parse(self, response):
        # 从meta中获取当前列表页的page对象
        page = response.meta["playwright_page"]
        # 等待所有可点击标题加载完成
        await page.wait_for_selector("h3.result-heading", timeout=10000)
        # 获取所有标题元素
        title_elements = await page.query_selector_all("h3.result-heading")

        for elem in title_elements:
            try:
                # 先监听新弹窗事件,再点击元素
                popup_promise = page.wait_for_event("popup", timeout=10000)
                await elem.click()
                # 获取弹出页对象
                popup_page = await popup_promise
                # 等待弹出页DOM加载完成
                await popup_page.wait_for_load_state("domcontentloaded", timeout=10000)
                
                # 提取弹出页数据
                body = await popup_page.locator("section#postingbody").all_text_contents()
                yield {
                    "body": [t.strip() for t in body if t.strip()]
                }

                # 关闭弹出页,释放资源
                await popup_page.close()
                # 可选:添加短延迟,避免触发反爬
                await page.wait_for_timeout(1000)
            except Exception as e:
                self.logger.error(f"处理条目出错: {str(e)}")
                continue
        
        # 全部处理完成后关闭列表页
        await page.close()

配置补充(settings.py)

确保项目settings.py中已添加scrapy-playwright必要配置:

# 启用playwright下载处理器
DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
# 指定异步 reactor
TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"
# 可选配置:设置浏览器启动参数,比如无头模式、代理等
PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": False, # 调试阶段可以设为False看浏览器操作过程
    "slow_mo": 500
}

注意事项

  • 如果点击后是当前页跳转而非新弹窗,不需要监听popup事件,点击后直接等待当前页加载,提取数据后调用page.go_back()返回列表页继续处理下一条即可
  • 可根据目标站点反爬策略调整等待时长、请求间隔,避免被拦截
  • 可以添加翻页逻辑,爬取多页列表的所有条目数据

内容的提问来源于stack exchange,提问作者an huynh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 17:36:00