基于Bull、Puppeteer、NodeJS和Express的爬虫任务队列改造咨询
Great question! Let's break down how to adapt your existing Express/Puppeteer setup to use Bull (Redis queues) and fix that Heroku timeout issue. The core problem here is your current request flow is synchronous—everything runs within a single web request cycle, which will hit Heroku's hard limit for long-running tasks.
Here's exactly how to restructure things, step by step:
1. First: Understand the New Flow
Instead of forcing the user to wait through the entire scrape/CSV process in one request, we'll:
- Accept the user's request, add the job to a Redis queue, and immediately return a job ID to the client
- Run the scraper and CSV generation asynchronously in the background via Bull
- Let the client check the job status later (or send them a notification when it's done)
Your current code structure is actually perfectly suited for this split—your existing createFile, getData, and buildCSVAndDeliver functions are modular and easy to move into a queue processor.
2. Setup Dependencies & Queue Instance
First, install the required packages:
npm install bull redis
Then create a dedicated queue file (e.g., queue.js) to manage your scrape jobs:
// queue.js const Queue = require('bull'); // Use Heroku's Redis addon URL, or local Redis for development const redisUrl = process.env.REDIS_URL || 'redis://localhost:6379'; const scrapeQueue = new Queue('scrape-jobs', redisUrl); module.exports = scrapeQueue;
Note: On Heroku, you'll need to add the Heroku Redis addon to your app to get the REDIS_URL environment variable.
3. Rewrite the POST Route
Instead of chaining your middleware, we'll now just add the job to the queue and respond immediately:
// Update your router file const scrapeQueue = require('./queue'); const Script = require("../models/Script"); router.post("/get-data/:id", ensureAuth, async (req, res) => { try { // Fetch the script once to pass its content to the queue const script = await Script.findById(req.params.id); // Add the job to the queue with all necessary data const job = await scrapeQueue.add({ scriptId: req.params.id, scriptContent: script.script, requestBody: req.body }); // Return a response to the client right away res.status(202).json({ message: "Scrape job started—check back for status", jobId: job.id // Client uses this to check progress }); } catch (error) { res.status(500).json({ error: "Failed to start scrape job. Please try again." }); } });
4. Move Your Scrape/CSV Logic to the Queue Processor
Take the logic from your three middleware functions and wrap them into a Bull job processor. Add this to your queue.js file:
// Add to queue.js const fs = require("fs"); const { Parser } = require("json2csv"); scrapeQueue.process(async (job) => { const { scriptId, scriptContent, requestBody } = job.data; let filename = `./scraper/${scriptId}.js`; let counter = 0; // Step 1: Create the script file (from createFile) while (fs.existsSync(filename)) { filename = `./scraper/${scriptId}${counter}.js`; counter++; } fs.writeFileSync(filename, scriptContent); try { // Step 2: Run the scraper (from getData) let scraper = require(`.${filename}`); const response = await scraper(requestBody); fs.unlinkSync(filename); // Clean up the script file // Step 3: Generate CSV (from buildCSVAndDeliver) // Important: Heroku's filesystem is temporary—store CSV in cloud storage (S3, etc.) instead! let csvFile = `./csv/${scriptId}.csv`; const parser = new Parser(); const csv = parser.parse(response); fs.writeFileSync(csvFile, csv); // Replace this with your delivery logic (e.g., upload to S3, send email to user) return { status: "success", csvPath: csvFile, // Or your cloud storage URL dataLength: response.length }; } catch (error) { // Clean up files on failure if (fs.existsSync(filename)) fs.unlinkSync(filename); throw new Error(`Scrape failed: ${error.message}`); } });
5. Add a Job Status Endpoint
Let clients check if their job is done and retrieve the CSV:
// Add to your router file const scrapeQueue = require('./queue'); router.get("/job-status/:jobId", ensureAuth, async (req, res) => { const job = await scrapeQueue.getJob(req.params.jobId); if (!job) { return res.status(404).json({ error: "Job not found" }); } const status = await job.getState(); const result = await job.returnvalue(); res.json({ jobId: job.id, status, // Can be 'waiting', 'active', 'completed', 'failed' result }); });
Key Adjustments to Note
- Heroku Filesystem Limitation: Heroku's local storage is wiped when your dyno restarts. Instead of saving CSVs locally, upload them to a cloud storage service like AWS S3 or Cloudinary, then return the public URL in the job result.
- Puppeteer on Heroku: Make sure you're using the official Puppeteer buildpack for Heroku to avoid missing Chrome dependencies.
- Error Handling: Bull automatically retries failed jobs by default—you can configure retry limits and delays in the queue setup if needed.
- Scalability: If you need more processing power, just spin up additional Heroku dynos—Bull will distribute jobs across all running processors automatically.
内容的提问来源于stack exchange,提问作者Riza Khan

