You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Puppeteer实现的Node.js爬虫API以提升抓取速度?

优化Node.js Puppeteer爬虫API性能方案

针对你基于Node.js+Puppeteer开发的Goodreads爬虫API请求耗时2-3秒的问题,以下是具体可落地的优化方案:

1. 复用Puppeteer浏览器实例(核心优化)

当前代码每次请求都启动和关闭浏览器,这是最大的性能开销。改为全局复用浏览器实例,仅在服务启动时初始化,关闭时销毁。

修改示例:

scraper-handler.ts

import { NextFunction, Request, Response } from "express";
import { MOST_POPULAR_LISTS } from "../utils/api/urls-endpoints.js";
import { listScraper } from "./spec-scrapers/list-scraper.js";
import puppeteer from "puppeteer";
import { GOODREADS_POPULAR_LISTS_URL } from "../utils/goodreads/urls.js";

// 全局复用浏览器实例
let browser: puppeteer.Browser | null = null;

// 服务启动时初始化浏览器
export const initBrowser = async () => {
  browser = await puppeteer.launch({
    headless: 'new', // 使用新版无头模式,性能更优
    args: [
      '--no-sandbox',
      '--disable-setuid-sandbox',
      '--disable-dev-shm-usage', // 避免内存不足问题
      '--disable-gpu',
      '--disable-images', // 直接禁用图片加载
    ],
  });
};

// 服务关闭时销毁浏览器
export const closeBrowser = async () => {
  if (browser) await browser.close();
};

export const scraperHandler = async (
  req: Request,
  res: Response,
  next: NextFunction
) => {
  if (!browser) {
    res.status(500).json({ status: "error", message: "Browser not initialized" });
    return;
  }

  // 每次请求创建新页面,而非复用默认页面
  const page = await browser.newPage();
  await page.setUserAgent(
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.102 Safari/537.36"
  );

  try {
    switch (req.url) {
      case `/${MOST_POPULAR_LISTS}`: {
        const result = await listScraper(
          page,
          GOODREADS_POPULAR_LISTS_URL,
          1,
          ".cell",
          ".listTitle",
          ".listTitle"
        );

        res.status(200).json({
          status: "success",
          data: result,
        });
        break;
      }
      default: {
        next();
        break;
      }
    }
  } finally {
    // 请求结束后关闭页面,释放资源
    await page.close();
  }
};

list-scraper.ts(对应修改)

import { Page } from "puppeteer";

export const listScraper = async (
  page: Page,
  url: string,
  pageI = 1,
  main: string,
  title = "",
  ref = ""
) => {
  await page.goto(url, {
    waitUntil: 'networkidle2', // 等待网络空闲(仅2个以下请求),比domcontentloaded更快
    timeout: 10000, // 设置超时时间
  });
  
  const books = await page.evaluate(
    (mainSelector, titleSelector, refSelector) => {
      const elements = document.querySelectorAll(mainSelector);
      
      return Array.from(elements)
        .slice(0, 3)
        .map((element) => {
          const title = element.querySelector(titleSelector)?.textContent?.trim();
          const ref = (element.querySelector(refSelector) as HTMLAnchorElement)?.href;

          return { title, ref };
        });
    },
    main,
    title,
    ref
  );

  return books;
};

2. 拦截不必要的资源加载

通过请求拦截,禁止加载图片、CSS、字体等非必要资源,进一步减少页面加载时间:

在页面创建后添加以下代码:

await page.setRequestInterception(true);
page.on('request', (req) => {
  const blockTypes = ['image', 'stylesheet', 'font', 'media'];
  if (blockTypes.includes(req.resourceType())) {
    req.abort();
  } else {
    req.continue();
  }
});

3. 替换为静态HTML解析(若页面无需JS渲染)

检查Goodreads目标页面:如果数据是直接渲染在HTML中的(而非通过AJAX加载),可以用axios+cheerio代替Puppeteer,速度提升数倍:

示例代码:

import axios from 'axios';
import cheerio from 'cheerio';

export const listScraper = async (url: string) => {
  const { data } = await axios.get(url, {
    headers: {
      'User-Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.102 Safari/537.36"
    }
  });

  const $ = cheerio.load(data);
  const books = [];

  $('.cell').slice(0,3).each((_, element) => {
    const title = $(element).find('.listTitle').text().trim();
    const ref = $(element).find('.listTitle').attr('href');
    books.push({ title, ref: `https://www.goodreads.com${ref}` });
  });

  return books;
};

4. 实现数据缓存

对抓取结果进行缓存,避免重复请求目标网站。可以用内存缓存(如node-cache)或Redis:

示例(node-cache):

import NodeCache from 'node-cache';
const cache = new NodeCache({ stdTTL: 300 }); // 缓存5分钟

export const scraperHandler = async (req: Request, res: Response, next: NextFunction) => {
  const cacheKey = `goodreads_popular_lists`;
  const cachedData = cache.get(cacheKey);

  if (cachedData) {
    return res.status(200).json({ status: "success", data: cachedData });
  }

  // 执行抓取逻辑
  const result = await listScraper(page, GOODREADS_POPULAR_LISTS_URL, 1, ".cell", ".listTitle", ".listTitle");
  
  cache.set(cacheKey, result);
  res.status(200).json({ status: "success", data: result });
};

5. 利用Node.js Cluster多核能力

Node.js默认单线程,用Cluster模块启动多个工作进程,充分利用CPU多核资源,提升并发处理能力:

主进程示例(server.ts):

import cluster from 'cluster';
import os from 'os';
import { initBrowser, closeBrowser } from './scraper-handler.js';

if (cluster.isPrimary) {
  const cpuCount = os.cpus().length;
  console.log(`主进程 ${process.pid} 启动,生成 ${cpuCount} 个工作进程`);

  for (let i = 0; i < cpuCount; i++) {
    cluster.fork();
  }

  cluster.on('exit', (worker) => {
    console.log(`工作进程 ${worker.process.pid} 退出,重启中...`);
    cluster.fork();
  });
} else {
  // 每个工作进程初始化自己的浏览器实例
  initBrowser().then(() => {
    // 启动Express服务
    import('./app.js').then(({ app }) => {
      app.listen(3000, () => {
        console.log(`工作进程 ${process.pid} 监听3000端口`);
      });
    });
  });

  process.on('exit', closeBrowser);
}

6. Puppeteer配置优化

  • 使用新版无头模式:headless: 'new'(比旧版headless: true性能更优)
  • 添加启动参数减少资源占用:--disable-dev-shm-usage、--no-sandbox、--disable-gpu等,如第一个优化点中的示例。

内容的提问来源于stack exchange,提问作者Ilia Popov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 03:14:52