You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Bull、Puppeteer、NodeJS和Express的爬虫任务队列改造咨询

Fixing Heroku 30-Second Timeout with Bull (Redis Queues) for Your Puppeteer Scraper

Great question! Let's break down how to adapt your existing Express/Puppeteer setup to use Bull (Redis queues) and fix that Heroku timeout issue. The core problem here is your current request flow is synchronous—everything runs within a single web request cycle, which will hit Heroku's hard limit for long-running tasks.

Here's exactly how to restructure things, step by step:


1. First: Understand the New Flow

Instead of forcing the user to wait through the entire scrape/CSV process in one request, we'll:

  • Accept the user's request, add the job to a Redis queue, and immediately return a job ID to the client
  • Run the scraper and CSV generation asynchronously in the background via Bull
  • Let the client check the job status later (or send them a notification when it's done)

Your current code structure is actually perfectly suited for this split—your existing createFile, getData, and buildCSVAndDeliver functions are modular and easy to move into a queue processor.


2. Setup Dependencies & Queue Instance

First, install the required packages:

npm install bull redis

Then create a dedicated queue file (e.g., queue.js) to manage your scrape jobs:

// queue.js
const Queue = require('bull');
// Use Heroku's Redis addon URL, or local Redis for development
const redisUrl = process.env.REDIS_URL || 'redis://localhost:6379';
const scrapeQueue = new Queue('scrape-jobs', redisUrl);

module.exports = scrapeQueue;

Note: On Heroku, you'll need to add the Heroku Redis addon to your app to get the REDIS_URL environment variable.


3. Rewrite the POST Route

Instead of chaining your middleware, we'll now just add the job to the queue and respond immediately:

// Update your router file
const scrapeQueue = require('./queue');
const Script = require("../models/Script");

router.post("/get-data/:id", ensureAuth, async (req, res) => {
  try {
    // Fetch the script once to pass its content to the queue
    const script = await Script.findById(req.params.id);
    
    // Add the job to the queue with all necessary data
    const job = await scrapeQueue.add({
      scriptId: req.params.id,
      scriptContent: script.script,
      requestBody: req.body
    });

    // Return a response to the client right away
    res.status(202).json({
      message: "Scrape job started—check back for status",
      jobId: job.id // Client uses this to check progress
    });
  } catch (error) {
    res.status(500).json({ error: "Failed to start scrape job. Please try again." });
  }
});

4. Move Your Scrape/CSV Logic to the Queue Processor

Take the logic from your three middleware functions and wrap them into a Bull job processor. Add this to your queue.js file:

// Add to queue.js
const fs = require("fs");
const { Parser } = require("json2csv");

scrapeQueue.process(async (job) => {
  const { scriptId, scriptContent, requestBody } = job.data;
  let filename = `./scraper/${scriptId}.js`;
  let counter = 0;

  // Step 1: Create the script file (from createFile)
  while (fs.existsSync(filename)) {
    filename = `./scraper/${scriptId}${counter}.js`;
    counter++;
  }
  fs.writeFileSync(filename, scriptContent);

  try {
    // Step 2: Run the scraper (from getData)
    let scraper = require(`.${filename}`);
    const response = await scraper(requestBody);
    fs.unlinkSync(filename); // Clean up the script file

    // Step 3: Generate CSV (from buildCSVAndDeliver)
    // Important: Heroku's filesystem is temporary—store CSV in cloud storage (S3, etc.) instead!
    let csvFile = `./csv/${scriptId}.csv`;
    const parser = new Parser();
    const csv = parser.parse(response);
    fs.writeFileSync(csvFile, csv);

    // Replace this with your delivery logic (e.g., upload to S3, send email to user)
    return {
      status: "success",
      csvPath: csvFile, // Or your cloud storage URL
      dataLength: response.length
    };
  } catch (error) {
    // Clean up files on failure
    if (fs.existsSync(filename)) fs.unlinkSync(filename);
    throw new Error(`Scrape failed: ${error.message}`);
  }
});

5. Add a Job Status Endpoint

Let clients check if their job is done and retrieve the CSV:

// Add to your router file
const scrapeQueue = require('./queue');

router.get("/job-status/:jobId", ensureAuth, async (req, res) => {
  const job = await scrapeQueue.getJob(req.params.jobId);
  
  if (!job) {
    return res.status(404).json({ error: "Job not found" });
  }

  const status = await job.getState();
  const result = await job.returnvalue();

  res.json({
    jobId: job.id,
    status, // Can be 'waiting', 'active', 'completed', 'failed'
    result
  });
});

Key Adjustments to Note

  1. Heroku Filesystem Limitation: Heroku's local storage is wiped when your dyno restarts. Instead of saving CSVs locally, upload them to a cloud storage service like AWS S3 or Cloudinary, then return the public URL in the job result.
  2. Puppeteer on Heroku: Make sure you're using the official Puppeteer buildpack for Heroku to avoid missing Chrome dependencies.
  3. Error Handling: Bull automatically retries failed jobs by default—you can configure retry limits and delays in the queue setup if needed.
  4. Scalability: If you need more processing power, just spin up additional Heroku dynos—Bull will distribute jobs across all running processors automatically.

内容的提问来源于stack exchange,提问作者Riza Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:07:44