You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用simplehtmldom抓取9000行大表格报Call to a member function find()错误

解决SimpleHTMLDom处理大表格时的"Call to a member function find() on boolean"错误

Hey Danny, let's work through this issue step by step. That error pops up because $html->find('table', 1) is returning false—either the second table doesn't exist in your full node_list.html file, or SimpleHTMLDom is failing to parse the large document entirely (even after adjusting memory limits). Here are actionable fixes to get your script working:

1. First, Verify the Target Table Actually Exists

Before diving into memory tweaks, confirm your full HTML file has at least two tables. Run this quick check to rule out file corruption or unexpected structure:

<?php
$content = file_get_contents('node_list.html');
// Count how many <table> tags are present in the file
$tableCount = substr_count($content, '<table');
echo "Total tables in file: " . $tableCount;
?>

If the count is less than 2, you'll need to adjust your find() index (maybe it should be 0 instead of 1). If it's 2 or more, move on to the next solutions.

2. Switch to PHP's Native DOM Extensions (Better for Large Files)

SimpleHTMLDom is great for small documents, but it’s memory-heavy for large datasets—loads the entire DOM into memory as a custom object. PHP’s built-in DOMDocument and DOMXPath are far more efficient with memory and handle big files better. Here’s a replacement script for your use case:

<?php
$dom = new DOMDocument();
// Suppress warnings about malformed HTML (common in real-world pages)
libxml_use_internal_errors(true);
$dom->loadHTMLFile('node_list.html');
libxml_clear_errors();

$xpath = new DOMXPath($dom);
// Target the second table (XPath uses 1-based indexing)
$table = $xpath->query('//table[2]')->item(0);

if (!$table) {
    die("Could not locate the second table");
}

$rowData = [];
foreach ($table->getElementsByTagName('tr') as $row) {
    $flight = [];
    foreach ($row->getElementsByTagName('td') as $cell) {
        $flight[] = trim($cell->textContent);
    }
    if (!empty($flight)) { // Skip empty rows to clean up data
        $rowData[] = $flight;
    }
}

print_r($rowData);
?>

This should handle your 9000-row table without hitting memory limits as quickly.

3. Bump the Memory Limit Higher

If you want to stick with SimpleHTMLDom, try increasing the memory limit further—128M might not be enough for a 9000-row table given SimpleHTMLDom’s memory-heavy structure. Run the script with:

php -d memory_limit=512M domtest.php

You can also track memory usage mid-script to see where it hits limits:

// Add this inside your loop to monitor memory usage
echo "Current memory usage: " . memory_get_usage(true) / 1024 / 1024 . "MB\n";

4. Parse Only the Target Table (Reduce Load)

Instead of loading the entire HTML file into SimpleHTMLDom, extract just the second table’s content first, then parse that smaller chunk. This cuts down the data SimpleHTMLDom has to process:

<?php
require('simplehtmldom/simple_html_dom.php');

$content = file_get_contents('node_list.html');
// Use regex to extract all tables (adjust if your HTML has complex table attributes)
preg_match_all('/<table[^>]*>(.*?)<\/table>/s', $content, $matches);

if (isset($matches[0][1])) { // Grab the second table (index 1)
    $html = str_get_html($matches[0][1]);
    $rowData = [];
    foreach($html->find('tr') as $row) {
        $flight = [];
        foreach($row->find('td') as $cell) {
            $flight[] = $cell->plaintext;
        }
        $rowData[] = $flight;
    }
    print_r($rowData);
} else {
    die("Second table not found in the file");
}
?>

Note: Regex isn’t perfect for HTML, but it works well for simple table structures like yours.

内容的提问来源于stack exchange,提问作者Danny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:43:41