PHP爬虫性能优化求助:8500条数据爬取速度待提升
Great question! Let's break down the key bottlenecks in your crawler and tackle them step by step—most of the slowdown is likely coming from network I/O (the biggest culprit) and inefficient string/DOM operations. Here are actionable optimizations tailored to your code:
1. Optimize Network Requests (Biggest Speed Win)
Right now, every call to getParts() creates a new stream context and makes a sequential HTTP request. Serial requests mean you're waiting for each page to load before moving to the next—this is where you'll get the most significant speed boost.
a. Reusable HTTP Context
You're recreating the stream context in both getSections() and getParts(). Move this to a static or class-level variable so you don't reinitialize it 54+ times:
// Define once at the top $httpContext = stream_context_create([ 'http' => [ 'method' => "GET", 'headers' => 'User-Agent: ACrawler/1.0\n' ] ]); // Reuse it in your functions @$doc->loadHTML(@file_get_contents($homePage, false, $httpContext));
b. Parallelize Requests
Instead of crawling one section at a time, use curl_multi to fetch multiple pages simultaneously. This overlaps network wait times. Here's a simplified adaptation of your getSections() to use parallel requests:
function getSections() { global $homePage, $httpContext, $sections; // ... (existing code to fetch and parse homepage links) $multiHandle = curl_multi_init(); $curlHandles = []; // Add all section links to curl multi foreach($result as $a){ $link = $homePage . $a->getAttribute("href"); $name = ""; // ... (existing name parsing code) $ch = curl_init($link); curl_setopt($ch, CURLOPT_USERAGENT, "ACrawler/1.0"); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_multi_add_handle($multiHandle, $ch); $curlHandles[(int)$ch] = ['name' => $name]; } // Execute parallel requests $running = null; do { curl_multi_exec($multiHandle, $running); curl_multi_select($multiHandle); } while ($running > 0); // Process responses foreach($curlHandles as $chId => $data){ $html = curl_multi_getcontent($chId); $doc = new DOMDocument(); @$doc->loadHTML($html); getPartsFromDoc($doc, $data['name']); // Modified getParts to accept a pre-loaded DOM curl_multi_remove_handle($multiHandle, $chId); } curl_multi_close($multiHandle); } // Modified getParts to avoid re-fetching pages function getPartsFromDoc($doc, $name) { global $sections; $newSection = ["sectionName" => $name, "parts" => []]; $xpath = new DOMXPath($doc); $result = $xpath->query("//a[@name]"); foreach($result as $a){ $newSection["parts"][] = newPart($a, $name); } $sections[] = $newSection; }
2. Optimize newPart() String Operations
Your newPart() function relies heavily on manual string splitting and looping—replacing these with regular expressions will reduce iteration count and speed up execution (plus make the code cleaner).
Example: Extract Number, Title, Categories with Regex
Instead of splitting the title into an array and looping through it, use a regex to capture all needed parts in one go:
function newPart($a, $name){ $newPart = []; $p = $a->getElementsByTagName("p")[0]; if (!$p) return $newPart; // Add null checks to avoid errors $bTag = $p->getElementsByTagName("b")[0]; if (!$bTag) return $newPart; $title = $bTag->textContent; // Regex to match course prefix, number, title, ints, and categories $pattern = '/^(\w+)\s+(\d+)\s+(.+?)(?:\s+\((\d+)\)\s+)?((?:\w+,?\s*)+)$/'; if (preg_match($pattern, $title, $matches)) { $newPart["number"] = $matches[2]; $newPart["title"] = trim($matches[3]); $newPart["ints"] = $matches[4] ?? null; // Split and clean categories $newPart["categories"] = array_map('trim', explode(',', $matches[5])); } // Optimize description parsing similarly with regex to replace loops $text = $p->textContent; $descriptionPattern = '/(.+?)(Prerequisite|Co-requisite.*)?View course details/'; if (preg_match($descriptionPattern, $text, $descMatches)) { $newPart["description"] = trim($descMatches[1]); $newPart["splitdesc"] = trim($descMatches[2] ?? ''); } $linkTag = $p->getElementsByTagName("a")[0]; $newPart["link"] = $linkTag?->getAttribute("href"); return $newPart; }
This replaces multiple loops with a single regex match, which is far faster for text parsing tasks—critical when executing 8500 times.
3. DOM & Variable Optimization
- Cache DOM Nodes: Avoid calling
getElementsByTagNamemultiple times for the same node. Cache results to reduce redundant DOM traversal:$pTags = $a->getElementsByTagName("p"); $p = $pTags->length > 0 ? $pTags[0] : null; if (!$p) return $newPart; $bTags = $p->getElementsByTagName("b"); $title = $bTags->length > 0 ? $bTags[0]->textContent : ''; - Ditch Global Variables: Using globals like
$homePageand$sectionsadds overhead and makes code harder to maintain. Wrap your crawler in a class instead:class CourseCrawler { private $homePage; private $httpContext; private $sections = []; public function __construct($homePage) { $this->homePage = $homePage; $this->httpContext = stream_context_create([ 'http' => [ 'method' => "GET", 'headers' => 'User-Agent: ACrawler/1.0\n' ] ]); } // Move all your functions here as methods, using $this-> instead of globals public function getSections() { /* ... */ } private function getPartsFromDoc($doc, $name) { /* ... */ } private function newPart($a, $name) { /* ... */ } } // Usage: $crawler = new CourseCrawler("google.com"); $crawler->getSections();
4. Disable DOMDocument Error Suppression (Optional)
The @ operator suppresses errors but adds small overhead. Instead, disable errors explicitly:
$doc = new DOMDocument(); $doc->loadHTML($html, LIBXML_NOWARNING | LIBXML_NOERROR);
This is cleaner and slightly faster than using @.
Final Notes
- Parallel Requests Will Give the Biggest Boost: Even a basic
curl_multiimplementation can cut your total runtime by 70-80% since you're no longer waiting for each page sequentially. - Regex Over Manual Splitting: Reduces loop count in
newPart(), which adds up drastically when executed 8500 times.
内容的提问来源于stack exchange,提问作者alexjvan

