You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP使用simple_html_dom爬取滚动加载数据的问题咨询

Hey there! Let's figure out how to grab all 11k items when your current setup with simple_html_dom only pulls the first 48. The key issue here is that simple_html_dom only parses the initial static HTML the server sends. Those extra batches of 48 items load dynamically via JavaScript as you scroll—something simple_html_dom can't handle on its own.

Here are two reliable ways to fix this:

Approach 1: Call the AJAX API Directly (Fastest & Cleanest)

Most sites with infinite scroll don't render new items directly in the page—they fetch raw data (usually JSON) via an AJAX endpoint. This is way more efficient than simulating scrolling because you can skip loading the entire page and just pull the data you need.

Here's how to do it:

  • Open your browser's DevTools (hit F12), go to the Network tab, and filter for "XHR" or "Fetch".
  • Scroll the target page slowly. You'll see new requests pop up that load the next 48 items. Check the request URL, parameters (like page, offset, limit, or a cursor value), and any required headers (like User-Agent or authorization tokens).
  • Once you've mapped out the endpoint and parameters, use PHP to loop through pages until you stop getting new data. You can use curl directly, or a library like Guzzle for easier handling.

Example with Guzzle:

require 'vendor/autoload.php';

$client = new GuzzleHttp\Client();
$allItems = [];
$page = 1;
$itemsPerPage = 48;

while (true) {
    try {
        // Replace with the actual AJAX endpoint and parameters
        $response = $client->get('https://example.com/api/items', [
            'query' => [
                'page' => $page,
                'limit' => $itemsPerPage
            ],
            'headers' => [
                'User-Agent' => 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
            ]
        ]);

        $data = json_decode($response->getBody(), true);
        
        // If no more items, break the loop
        if (empty($data['items'])) {
            break;
        }

        $allItems = array_merge($allItems, $data['items']);
        $page++;

        // Add a delay to avoid overwhelming the server
        sleep(1);
    } catch (Exception $e) {
        echo "Error fetching page $page: " . $e->getMessage() . "\n";
        break;
    }
}

// Now $allItems has all 11k+ items
print_r(count($allItems) . " items fetched!");
Approach 2: Simulate Scrolling with a Headless Browser (For Tricky Sites)

If the site hides its AJAX endpoint behind complex logic or uses anti-scraping measures, you can simulate a real user scrolling using a headless browser. For PHP, spatie/browsershot is a great tool—it wraps Puppeteer (a Node.js tool) to control a headless Chrome/Chromium instance.

Steps:

  1. Install Browsershot via Composer:
    composer require spatie/browsershot
    
  2. Make sure you have Node.js and Puppeteer installed (Browsershot will prompt you to install Puppeteer if it's missing).
  3. Write code to navigate to the page, scroll repeatedly until no new content loads, then grab the full HTML to parse with simple_html_dom.

Example Code:

require 'vendor/autoload.php';
require 'simple_html_dom.php'; // Include your simple_html_dom file

use Spatie\Browsershot\Browsershot;

// Navigate to the page and scroll to load all content
$html = Browsershot::url('https://example.com/your-target-page')
    ->setUserAgent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36')
    ->scrollToBottom()
    ->waitUntilNetworkIdle() // Wait for all AJAX requests to finish
    ->bodyHtml();

// Now parse the full HTML with simple_html_dom
$htmlDom = str_get_html($html);
$items = $htmlDom->find('.item-class'); // Replace with your item selector

echo count($items) . " items found!\n";

// Process each item as needed
foreach ($items as $item) {
    echo $item->plaintext . "\n";
}

Important Notes:

  • Always respect the site's robots.txt and terms of service. Don't send too many requests too quickly—add delays (like sleep(1)) between calls to avoid getting blocked.
  • Some sites might require you to set a valid User-Agent header to mimic a real browser, otherwise they might block your requests or return empty data.
  • If you run into CAPTCHAs or rate limits, you might need to use proxies or rotate IP addresses (but only if allowed by the site's rules).

内容的提问来源于stack exchange,提问作者rufatZZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:41:08