You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于PHP Simple HTML DOM Parser优化多网站批量爬取效率?

Absolutely! Splitting your scraping logic into separate PHP files is a fantastic way to clean up your code, make it easier to maintain, and lay the groundwork for actually boosting your scraping efficiency. Let’s break down how to do this, plus some key tweaks to speed things up beyond just code organization.

1. Modularize Your Scraping Logic

First, split out site-specific parsing rules into dedicated files. This way, if a target site changes its HTML structure, you only need to update one file instead of digging through a monolithic script.

For example:

  • Create ScraperSiteA.php for your first target site:
<?php
require_once('simple_html_dom.php');

class ScraperSiteA {
    public function scrape(string $url): array {
        $html = file_get_html($url);
        if (!$html) return [];
        
        $collection = $html->find('div.info');
        $results = [];
        
        foreach ($collection as $item) {
            // Custom parsing for Site A
            $results[] = [
                'title' => $item->find('h2.title', 0)->plaintext ?? 'No title',
                'details' => $item->find('p.description', 0)->plaintext ?? ''
            ];
        }
        
        $html->clear(); // Critical to free up memory!
        return $results;
    }
}
?>
  • Repeat this for other sites (e.g., ScraperSiteB.php) with their own selectors.

Then, make a main orchestrator file (like BatchScraper.php) to manage all scrapers:

<?php
require_once('ScraperSiteA.php');
require_once('ScraperSiteB.php');

// Group URLs by their target site
$targets = [
    'siteA' => [
        'https://sitea.com/latest',
        'https://sitea.com/archive'
    ],
    'siteB' => [
        'https://siteb.com/category/news',
        'https://siteb.com/category/updates'
    ]
];

// Initialize scrapers
$scrapers = [
    'siteA' => new ScraperSiteA(),
    'siteB' => new ScraperSiteB()
];

$allResults = [];

// Sequential scrape (good for starting out)
foreach ($targets as $site => $urls) {
    foreach ($urls as $url) {
        echo "Scraping $url...\n";
        $allResults[$site][] = $scrapers[$site]->scrape($url);
    }
}

// Do something with your results (save to DB, export to CSV, etc.)
print_r($allResults);
?>

2. Boost Efficiency with Parallel Processing

The biggest win for speed comes from moving away from sequential scraping (one page at a time) to parallel requests. Since you’ve already modularized your code, adding parallelism is straightforward using PHP’s curl_multi extension.

Here’s a simplified example of parallel scraping:

<?php
require_once('simple_html_dom.php');

// Reusable scraping function (can be moved to a shared file too)
function parseHtml(string $htmlContent): array {
    $html = str_get_html($htmlContent);
    if (!$html) return [];
    
    $collection = $html->find('div.info');
    $results = [];
    
    foreach ($collection as $item) {
        $results[] = [
            'title' => $item->find('h2', 0)->plaintext ?? '',
            'content' => $item->find('p', 0)->plaintext ?? ''
        ];
    }
    
    $html->clear();
    return $results;
}

// List of all URLs to scrape
$urls = [
    'https://sitea.com/latest',
    'https://sitea.com/archive',
    'https://siteb.com/category/news',
    'https://siteb.com/category/updates'
];

// Set up multi-curl
$multiHandle = curl_multi_init();
$curlHandles = [];

foreach ($urls as $url) {
    $ch = curl_init($url);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true);
    curl_setopt($ch, CURLOPT_USERAGENT, 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36');
    
    $curlHandles[$url] = $ch;
    curl_multi_add_handle($multiHandle, $ch);
}

// Execute all requests in parallel
$running = null;
do {
    curl_multi_exec($multiHandle, $running);
    curl_multi_select($multiHandle);
} while ($running > 0);

// Process results
$allResults = [];
foreach ($curlHandles as $url => $ch) {
    $htmlContent = curl_multi_getcontent($ch);
    if ($htmlContent) {
        $allResults[$url] = parseHtml($htmlContent);
    }
    
    curl_multi_remove_handle($multiHandle, $ch);
    curl_close($ch);
}

curl_multi_close($multiHandle);

print_r($allResults);
?>

3. Extra Tips for Smooth, Fast Scraping

  • Respect Rate Limits: Add small delays (1-2 seconds) between requests if you’re not using parallelism, and check each site’s robots.txt to avoid getting blocked.
  • Cache Responses: Store scraped HTML locally (e.g., in a file or database) for a few hours/days so you don’t re-scrape the same pages unnecessarily.
  • Memory Management: Never skip $html->clear()—it frees up memory, which is crucial when scraping dozens of pages.
  • Handle Errors: Add try/catch blocks and check for failed requests (e.g., file_get_html returning false) to avoid crashing your script.

Splitting into modular files makes your code way easier to maintain, and pairing that with parallel processing will cut down your scraping time drastically. Start small with the modular setup, then add parallelism once you’re comfortable!

内容的提问来源于stack exchange,提问作者Mo Mononoke

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:24:32