如何扩展Crawlee的enqueueLinks函数添加自定义入队逻辑?
问题:扩展Crawlee的enqueueLinks实现自定义逻辑
问题背景
我正在使用crawlee@3.0.4,按照官方快速教程搭建了如下PlaywrightCrawler爬虫:
import { PlaywrightCrawler } from 'crawlee'; const crawler = new PlaywrightCrawler({ requestHandler: async ({ page, request, enqueueLinks }) => { console.log(`Processing: ${request.url}`) await page.waitForSelector('.ActorStorePagination-pages a'); await enqueueLinks({ selector: '.ActorStorePagination-pages > a', label: 'LIST', }) }, });
现在我需要扩展传入requestHandler的enqueueLinks函数,实现每当有新URL添加到队列时执行自定义逻辑,例如统计特定类型链接的数量,以便进行额外日志记录或向其他服务推送消息。请问是否有可行的实现方式?
我曾尝试继承PlaywrightCrawler类,但由于requestHandler被包裹在对象中,无法访问自定义类的属性,代码如下:
class CustomCrawler extends PlaywrightCrawler { categoryPagesQueued: string[]; constructor() { super({ requestHandler: async ({ page, request, enqueueLinks }) => { console.log(`Processing: ${request.url}`) // Wait for the actor cards to render, // otherwise enqueueLinks wouldn't enqueue anything. await page.waitForSelector('.ActorStorePagination-pages a'); // Error: this does not access the CustomCrawler.categoryPagesQueued this.categoryPagesQueued.push("foo"); customLogic(this.categoryPagesQueued); await enqueueLinks({ selector: '.ActorStorePagination-pages > a', label: 'LIST', }) }, }) } }
解决方案
方案1:利用enqueueLinks的transformRequestFunction钩子
enqueueLinks内置了transformRequestFunction参数,可直接在该函数中对即将入队的请求做处理,同时嵌入自定义逻辑:
import { PlaywrightCrawler } from 'crawlee'; // 统计变量,可根据需求存入全局存储或外部数据库 let listLinkCount = 0; const crawler = new PlaywrightCrawler({ requestHandler: async ({ page, request, enqueueLinks }) => { console.log(`Processing: ${request.url}`); await page.waitForSelector('.ActorStorePagination-pages a'); await enqueueLinks({ selector: '.ActorStorePagination-pages > a', label: 'LIST', transformRequestFunction: (request) => { // 执行自定义统计逻辑 if (request.label === 'LIST') { listLinkCount++; console.log(`已入队LIST类型链接数量:${listLinkCount}`); // 此处可添加消息推送、日志上报等逻辑 } // 返回原请求(也可按需修改请求属性) return request; } }); }, });
方案2:正确绑定this到自定义爬虫类
如果需要维护爬虫实例的内部状态,可将requestHandler抽为类方法并绑定this,解决指向问题:
import { PlaywrightCrawler } from 'crawlee'; class CustomCrawler extends PlaywrightCrawler { categoryPagesQueued: string[] = []; constructor() { super({ // 绑定类方法的this指向当前实例 requestHandler: this.handleRequest.bind(this), }); } // 抽离为类方法,确保this指向当前CustomCrawler实例 async handleRequest({ page, request, enqueueLinks }) { console.log(`Processing: ${request.url}`); await page.waitForSelector('.ActorStorePagination-pages a'); // 现在可正常访问类的属性 this.categoryPagesQueued.push(request.url); console.log(`已入队分类页面数量:${this.categoryPagesQueued.length}`); await enqueueLinks({ selector: '.ActorStorePagination-pages > a', label: 'LIST', // 结合transformRequestFunction做细粒度处理 transformRequestFunction: (req) => { this.categoryPagesQueued.push(req.url); return req; } }); } }
方案3:监听RequestQueue的add事件
Crawlee的请求队列支持事件监听,可全局捕获所有入队请求,覆盖所有入队场景:
import { PlaywrightCrawler, RequestQueue } from 'crawlee'; (async () => { const requestQueue = await RequestQueue.open(); // 监听队列的add事件,所有入队请求都会触发 requestQueue.on('add', (request) => { if (request.label === 'LIST') { console.log(`新入队LIST链接:${request.url}`); // 执行统计、推送等自定义逻辑 } }); const crawler = new PlaywrightCrawler({ requestQueue, requestHandler: async ({ page, request, enqueueLinks }) => { console.log(`Processing: ${request.url}`); await page.waitForSelector('.ActorStorePagination-pages a'); await enqueueLinks({ selector: '.ActorStorePagination-pages > a', label: 'LIST', }); }, }); await crawler.run(['你的起始URL']); })();
方案选型说明
- 方案1:最轻量化,仅针对
enqueueLinks触发的入队请求处理,适合简单统计场景; - 方案2:适合需要维护爬虫实例内部状态的复杂场景;
- 方案3:全局监听队列,覆盖所有入队方式(包括手动入队),适用范围最广。
内容的提问来源于stack exchange,提问作者toanphan19
相关产品推荐
相关产品推荐

