如何优化Puppeteer实现的Node.js爬虫API以提升抓取速度?
优化Node.js Puppeteer爬虫API性能方案
针对你基于Node.js+Puppeteer开发的Goodreads爬虫API请求耗时2-3秒的问题,以下是具体可落地的优化方案:
1. 复用Puppeteer浏览器实例(核心优化)
当前代码每次请求都启动和关闭浏览器,这是最大的性能开销。改为全局复用浏览器实例,仅在服务启动时初始化,关闭时销毁。
修改示例:
scraper-handler.ts
import { NextFunction, Request, Response } from "express"; import { MOST_POPULAR_LISTS } from "../utils/api/urls-endpoints.js"; import { listScraper } from "./spec-scrapers/list-scraper.js"; import puppeteer from "puppeteer"; import { GOODREADS_POPULAR_LISTS_URL } from "../utils/goodreads/urls.js"; // 全局复用浏览器实例 let browser: puppeteer.Browser | null = null; // 服务启动时初始化浏览器 export const initBrowser = async () => { browser = await puppeteer.launch({ headless: 'new', // 使用新版无头模式,性能更优 args: [ '--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage', // 避免内存不足问题 '--disable-gpu', '--disable-images', // 直接禁用图片加载 ], }); }; // 服务关闭时销毁浏览器 export const closeBrowser = async () => { if (browser) await browser.close(); }; export const scraperHandler = async ( req: Request, res: Response, next: NextFunction ) => { if (!browser) { res.status(500).json({ status: "error", message: "Browser not initialized" }); return; } // 每次请求创建新页面,而非复用默认页面 const page = await browser.newPage(); await page.setUserAgent( "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.102 Safari/537.36" ); try { switch (req.url) { case `/${MOST_POPULAR_LISTS}`: { const result = await listScraper( page, GOODREADS_POPULAR_LISTS_URL, 1, ".cell", ".listTitle", ".listTitle" ); res.status(200).json({ status: "success", data: result, }); break; } default: { next(); break; } } } finally { // 请求结束后关闭页面,释放资源 await page.close(); } };
list-scraper.ts(对应修改)
import { Page } from "puppeteer"; export const listScraper = async ( page: Page, url: string, pageI = 1, main: string, title = "", ref = "" ) => { await page.goto(url, { waitUntil: 'networkidle2', // 等待网络空闲(仅2个以下请求),比domcontentloaded更快 timeout: 10000, // 设置超时时间 }); const books = await page.evaluate( (mainSelector, titleSelector, refSelector) => { const elements = document.querySelectorAll(mainSelector); return Array.from(elements) .slice(0, 3) .map((element) => { const title = element.querySelector(titleSelector)?.textContent?.trim(); const ref = (element.querySelector(refSelector) as HTMLAnchorElement)?.href; return { title, ref }; }); }, main, title, ref ); return books; };
2. 拦截不必要的资源加载
通过请求拦截,禁止加载图片、CSS、字体等非必要资源,进一步减少页面加载时间:
在页面创建后添加以下代码:
await page.setRequestInterception(true); page.on('request', (req) => { const blockTypes = ['image', 'stylesheet', 'font', 'media']; if (blockTypes.includes(req.resourceType())) { req.abort(); } else { req.continue(); } });
3. 替换为静态HTML解析(若页面无需JS渲染)
检查Goodreads目标页面:如果数据是直接渲染在HTML中的(而非通过AJAX加载),可以用axios+cheerio代替Puppeteer,速度提升数倍:
示例代码:
import axios from 'axios'; import cheerio from 'cheerio'; export const listScraper = async (url: string) => { const { data } = await axios.get(url, { headers: { 'User-Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.102 Safari/537.36" } }); const $ = cheerio.load(data); const books = []; $('.cell').slice(0,3).each((_, element) => { const title = $(element).find('.listTitle').text().trim(); const ref = $(element).find('.listTitle').attr('href'); books.push({ title, ref: `https://www.goodreads.com${ref}` }); }); return books; };
4. 实现数据缓存
对抓取结果进行缓存,避免重复请求目标网站。可以用内存缓存(如node-cache)或Redis:
示例(node-cache):
import NodeCache from 'node-cache'; const cache = new NodeCache({ stdTTL: 300 }); // 缓存5分钟 export const scraperHandler = async (req: Request, res: Response, next: NextFunction) => { const cacheKey = `goodreads_popular_lists`; const cachedData = cache.get(cacheKey); if (cachedData) { return res.status(200).json({ status: "success", data: cachedData }); } // 执行抓取逻辑 const result = await listScraper(page, GOODREADS_POPULAR_LISTS_URL, 1, ".cell", ".listTitle", ".listTitle"); cache.set(cacheKey, result); res.status(200).json({ status: "success", data: result }); };
5. 利用Node.js Cluster多核能力
Node.js默认单线程,用Cluster模块启动多个工作进程,充分利用CPU多核资源,提升并发处理能力:
主进程示例(server.ts):
import cluster from 'cluster'; import os from 'os'; import { initBrowser, closeBrowser } from './scraper-handler.js'; if (cluster.isPrimary) { const cpuCount = os.cpus().length; console.log(`主进程 ${process.pid} 启动,生成 ${cpuCount} 个工作进程`); for (let i = 0; i < cpuCount; i++) { cluster.fork(); } cluster.on('exit', (worker) => { console.log(`工作进程 ${worker.process.pid} 退出,重启中...`); cluster.fork(); }); } else { // 每个工作进程初始化自己的浏览器实例 initBrowser().then(() => { // 启动Express服务 import('./app.js').then(({ app }) => { app.listen(3000, () => { console.log(`工作进程 ${process.pid} 监听3000端口`); }); }); }); process.on('exit', closeBrowser); }
6. Puppeteer配置优化
- 使用新版无头模式:
headless: 'new'(比旧版headless: true性能更优) - 添加启动参数减少资源占用:
--disable-dev-shm-usage、--no-sandbox、--disable-gpu等,如第一个优化点中的示例。
内容的提问来源于stack exchange,提问作者Ilia Popov
相关产品推荐
相关产品推荐

