You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP爬虫性能优化求助:8500条数据爬取速度待提升

Great question! Let's break down the key bottlenecks in your crawler and tackle them step by step—most of the slowdown is likely coming from network I/O (the biggest culprit) and inefficient string/DOM operations. Here are actionable optimizations tailored to your code:

1. Optimize Network Requests (Biggest Speed Win)

Right now, every call to getParts() creates a new stream context and makes a sequential HTTP request. Serial requests mean you're waiting for each page to load before moving to the next—this is where you'll get the most significant speed boost.

a. Reusable HTTP Context

You're recreating the stream context in both getSections() and getParts(). Move this to a static or class-level variable so you don't reinitialize it 54+ times:

// Define once at the top
$httpContext = stream_context_create([
    'http' => [
        'method' => "GET",
        'headers' => 'User-Agent: ACrawler/1.0\n'
    ]
]);

// Reuse it in your functions
@$doc->loadHTML(@file_get_contents($homePage, false, $httpContext));

b. Parallelize Requests

Instead of crawling one section at a time, use curl_multi to fetch multiple pages simultaneously. This overlaps network wait times. Here's a simplified adaptation of your getSections() to use parallel requests:

function getSections() {
    global $homePage, $httpContext, $sections;
    // ... (existing code to fetch and parse homepage links)

    $multiHandle = curl_multi_init();
    $curlHandles = [];

    // Add all section links to curl multi
    foreach($result as $a){
        $link = $homePage . $a->getAttribute("href");
        $name = "";
        // ... (existing name parsing code)
        
        $ch = curl_init($link);
        curl_setopt($ch, CURLOPT_USERAGENT, "ACrawler/1.0");
        curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
        curl_multi_add_handle($multiHandle, $ch);
        $curlHandles[(int)$ch] = ['name' => $name];
    }

    // Execute parallel requests
    $running = null;
    do {
        curl_multi_exec($multiHandle, $running);
        curl_multi_select($multiHandle);
    } while ($running > 0);

    // Process responses
    foreach($curlHandles as $chId => $data){
        $html = curl_multi_getcontent($chId);
        $doc = new DOMDocument();
        @$doc->loadHTML($html);
        getPartsFromDoc($doc, $data['name']); // Modified getParts to accept a pre-loaded DOM
        curl_multi_remove_handle($multiHandle, $chId);
    }
    curl_multi_close($multiHandle);
}

// Modified getParts to avoid re-fetching pages
function getPartsFromDoc($doc, $name) {
    global $sections;
    $newSection = ["sectionName" => $name, "parts" => []];
    $xpath = new DOMXPath($doc);
    $result = $xpath->query("//a[@name]");
    
    foreach($result as $a){
        $newSection["parts"][] = newPart($a, $name);
    }
    $sections[] = $newSection;
}

2. Optimize newPart() String Operations

Your newPart() function relies heavily on manual string splitting and looping—replacing these with regular expressions will reduce iteration count and speed up execution (plus make the code cleaner).

Example: Extract Number, Title, Categories with Regex

Instead of splitting the title into an array and looping through it, use a regex to capture all needed parts in one go:

function newPart($a, $name){
    $newPart = [];
    $p = $a->getElementsByTagName("p")[0];
    if (!$p) return $newPart; // Add null checks to avoid errors
    
    $bTag = $p->getElementsByTagName("b")[0];
    if (!$bTag) return $newPart;
    $title = $bTag->textContent;

    // Regex to match course prefix, number, title, ints, and categories
    $pattern = '/^(\w+)\s+(\d+)\s+(.+?)(?:\s+\((\d+)\)\s+)?((?:\w+,?\s*)+)$/';
    if (preg_match($pattern, $title, $matches)) {
        $newPart["number"] = $matches[2];
        $newPart["title"] = trim($matches[3]);
        $newPart["ints"] = $matches[4] ?? null;
        // Split and clean categories
        $newPart["categories"] = array_map('trim', explode(',', $matches[5]));
    }

    // Optimize description parsing similarly with regex to replace loops
    $text = $p->textContent;
    $descriptionPattern = '/(.+?)(Prerequisite|Co-requisite.*)?View course details/';
    if (preg_match($descriptionPattern, $text, $descMatches)) {
        $newPart["description"] = trim($descMatches[1]);
        $newPart["splitdesc"] = trim($descMatches[2] ?? '');
    }

    $linkTag = $p->getElementsByTagName("a")[0];
    $newPart["link"] = $linkTag?->getAttribute("href");
    
    return $newPart;
}

This replaces multiple loops with a single regex match, which is far faster for text parsing tasks—critical when executing 8500 times.

3. DOM & Variable Optimization

  • Cache DOM Nodes: Avoid calling getElementsByTagName multiple times for the same node. Cache results to reduce redundant DOM traversal:
    $pTags = $a->getElementsByTagName("p");
    $p = $pTags->length > 0 ? $pTags[0] : null;
    if (!$p) return $newPart;
    
    $bTags = $p->getElementsByTagName("b");
    $title = $bTags->length > 0 ? $bTags[0]->textContent : '';
    
  • Ditch Global Variables: Using globals like $homePage and $sections adds overhead and makes code harder to maintain. Wrap your crawler in a class instead:
    class CourseCrawler {
        private $homePage;
        private $httpContext;
        private $sections = [];
    
        public function __construct($homePage) {
            $this->homePage = $homePage;
            $this->httpContext = stream_context_create([
                'http' => [
                    'method' => "GET",
                    'headers' => 'User-Agent: ACrawler/1.0\n'
                ]
            ]);
        }
    
        // Move all your functions here as methods, using $this-> instead of globals
        public function getSections() { /* ... */ }
        private function getPartsFromDoc($doc, $name) { /* ... */ }
        private function newPart($a, $name) { /* ... */ }
    }
    
    // Usage:
    $crawler = new CourseCrawler("google.com");
    $crawler->getSections();
    

4. Disable DOMDocument Error Suppression (Optional)

The @ operator suppresses errors but adds small overhead. Instead, disable errors explicitly:

$doc = new DOMDocument();
$doc->loadHTML($html, LIBXML_NOWARNING | LIBXML_NOERROR);

This is cleaner and slightly faster than using @.

Final Notes

  • Parallel Requests Will Give the Biggest Boost: Even a basic curl_multi implementation can cut your total runtime by 70-80% since you're no longer waiting for each page sequentially.
  • Regex Over Manual Splitting: Reduces loop count in newPart(), which adds up drastically when executed 8500 times.

内容的提问来源于stack exchange,提问作者alexjvan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:30:38