You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否实现Chromium常驻实例,通过Puppeteer/PHP提交采集任务?

Answers to Your Puppeteer & PHP Chromium Questions

Great questions! Let's break this down step by step—both of your goals are totally achievable, and I'll walk you through how to approach each one.

1. Can I keep a Chromium instance running permanently with Puppeteer, reusing it for multiple scraping tasks?

Absolutely! This is actually a performance best practice, since launching a new Chromium instance every time adds significant overhead (time, memory, CPU). Here's how to pull it off:

  • Initialize a single Browser instance once when your application starts up.
  • For each scraping task, create a new Page within this existing browser (no need to spin up a whole new browser).
  • After finishing a task, close the individual Page (not the entire Browser) to free up resources while keeping the browser alive for future work.

Example: Node.js Puppeteer Service with Persistent Browser

const puppeteer = require('puppeteer');
const express = require('express');
const app = express();
app.use(express.json());

// Launch Chromium once at startup
let browser;
(async () => {
  browser = await puppeteer.launch({
    headless: 'new', // Use the modern headless mode for better efficiency
    args: ['--no-sandbox', '--disable-setuid-sandbox'] // Adjust based on your server environment
  });
  console.log('Chromium instance is up and ready');
})();

// Endpoint to handle incoming scraping tasks
app.post('/scrape', async (req, res) => {
  try {
    const { url } = req.body;
    if (!browser) {
      return res.status(500).json({ error: 'Browser still initializing—try again in a moment' });
    }

    // Create a fresh page for this task
    const page = await browser.newPage();
    await page.goto(url, { waitUntil: 'networkidle2' });
    const pageTitle = await page.title();
    const pageContent = await page.content();

    // Clean up the page (keep the browser running!)
    await page.close();

    res.json({ 
      url, 
      title: pageTitle, 
      content_preview: pageContent.slice(0, 500) // Trim content for brevity
    });
  } catch (error) {
    res.status(500).json({ error: error.message });
  }
});

app.listen(3000, () => {
  console.log('Scraping service running on port 3000');
});

This setup keeps Chromium running in the background, and every /scrape request uses a fresh page within the same browser instance.

2. Can I implement a persistent Chromium instance and send requests via PHP?

PHP is a request-driven language—by default, PHP processes terminate after each web request, so it can't directly maintain a persistent Chromium instance on its own. But you can make this work with a two-part setup:

  • A permanent backend service (using Node.js, like the example above) that manages the Chromium instance.
  • PHP as a client that sends HTTP requests to this backend to trigger scraping tasks.

Example: PHP Client Calling the Puppeteer Service

<?php
// PHP script to send a scraping task to the Node.js service
$targetUrl = 'https://example.com';
$serviceEndpoint = 'http://localhost:3000/scrape';

// Prepare request data
$requestData = json_encode(['url' => $targetUrl]);

// Initialize cURL
$curlHandle = curl_init($serviceEndpoint);
curl_setopt($curlHandle, CURLOPT_POST, true);
curl_setopt($curlHandle, CURLOPT_POSTFIELDS, $requestData);
curl_setopt($curlHandle, CURLOPT_HTTPHEADER, ['Content-Type: application/json']);
curl_setopt($curlHandle, CURLOPT_RETURNTRANSFER, true);

// Execute the request
$response = curl_exec($curlHandle);
$httpStatus = curl_getinfo($curlHandle, CURLINFO_HTTP_CODE);
curl_close($curlHandle);

// Process the response
if ($httpStatus === 200) {
    $result = json_decode($response, true);
    echo "Scraped Title: " . $result['title'] . "\n";
    echo "Content Preview: " . $result['content_preview'] . "\n";
} else {
    echo "Request failed: " . $response . "\n";
}
?>

Keeping the Node.js Service Running Permanently

To make sure the Node.js service (and its Chromium instance) stays up even after server restarts, use a process manager like pm2:

  • Install PM2 globally: npm install -g pm2
  • Start your service: pm2 start your-scraper-service.js
  • PM2 will auto-restart the service if it crashes and keep it running in the background.

内容的提问来源于stack exchange,提问作者Imobach

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:44:43