PHP使用simple_html_dom爬取滚动加载数据的问题咨询
Hey there! Let's figure out how to grab all 11k items when your current setup with simple_html_dom only pulls the first 48. The key issue here is that simple_html_dom only parses the initial static HTML the server sends. Those extra batches of 48 items load dynamically via JavaScript as you scroll—something simple_html_dom can't handle on its own.
Here are two reliable ways to fix this:
Most sites with infinite scroll don't render new items directly in the page—they fetch raw data (usually JSON) via an AJAX endpoint. This is way more efficient than simulating scrolling because you can skip loading the entire page and just pull the data you need.
Here's how to do it:
- Open your browser's DevTools (hit F12), go to the Network tab, and filter for "XHR" or "Fetch".
- Scroll the target page slowly. You'll see new requests pop up that load the next 48 items. Check the request URL, parameters (like
page,offset,limit, or acursorvalue), and any required headers (likeUser-Agentor authorization tokens). - Once you've mapped out the endpoint and parameters, use PHP to loop through pages until you stop getting new data. You can use
curldirectly, or a library like Guzzle for easier handling.
Example with Guzzle:
require 'vendor/autoload.php'; $client = new GuzzleHttp\Client(); $allItems = []; $page = 1; $itemsPerPage = 48; while (true) { try { // Replace with the actual AJAX endpoint and parameters $response = $client->get('https://example.com/api/items', [ 'query' => [ 'page' => $page, 'limit' => $itemsPerPage ], 'headers' => [ 'User-Agent' => 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' ] ]); $data = json_decode($response->getBody(), true); // If no more items, break the loop if (empty($data['items'])) { break; } $allItems = array_merge($allItems, $data['items']); $page++; // Add a delay to avoid overwhelming the server sleep(1); } catch (Exception $e) { echo "Error fetching page $page: " . $e->getMessage() . "\n"; break; } } // Now $allItems has all 11k+ items print_r(count($allItems) . " items fetched!");
If the site hides its AJAX endpoint behind complex logic or uses anti-scraping measures, you can simulate a real user scrolling using a headless browser. For PHP, spatie/browsershot is a great tool—it wraps Puppeteer (a Node.js tool) to control a headless Chrome/Chromium instance.
Steps:
- Install Browsershot via Composer:
composer require spatie/browsershot - Make sure you have Node.js and Puppeteer installed (Browsershot will prompt you to install Puppeteer if it's missing).
- Write code to navigate to the page, scroll repeatedly until no new content loads, then grab the full HTML to parse with simple_html_dom.
Example Code:
require 'vendor/autoload.php'; require 'simple_html_dom.php'; // Include your simple_html_dom file use Spatie\Browsershot\Browsershot; // Navigate to the page and scroll to load all content $html = Browsershot::url('https://example.com/your-target-page') ->setUserAgent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36') ->scrollToBottom() ->waitUntilNetworkIdle() // Wait for all AJAX requests to finish ->bodyHtml(); // Now parse the full HTML with simple_html_dom $htmlDom = str_get_html($html); $items = $htmlDom->find('.item-class'); // Replace with your item selector echo count($items) . " items found!\n"; // Process each item as needed foreach ($items as $item) { echo $item->plaintext . "\n"; }
Important Notes:
- Always respect the site's
robots.txtand terms of service. Don't send too many requests too quickly—add delays (likesleep(1)) between calls to avoid getting blocked. - Some sites might require you to set a valid
User-Agentheader to mimic a real browser, otherwise they might block your requests or return empty data. - If you run into CAPTCHAs or rate limits, you might need to use proxies or rotate IP addresses (but only if allowed by the site's rules).
内容的提问来源于stack exchange,提问作者rufatZZ

