使用html_dom爬取网站遇file_get_html超时及流打开失败问题求助
Hey there, I’ve run into this exact issue multiple times while scraping with simple_html_dom—network lag can turn a quick scrape into a frustrating timeout error. Let’s break down the best fixes to get this under control:
1. Add Custom Timeout Context to file_get_html
The file_get_html function actually accepts a stream context parameter that lets you override default network timeouts. This ensures your request doesn’t hang for 30+ seconds waiting for a slow server.
Here’s how to implement it:
// Create a stream context with strict timeouts $context = stream_context_create([ 'http' => [ 'connect_timeout' => 5, // Max time to establish a connection (seconds) 'read_timeout' => 5, // Max time to wait for data after connecting 'timeout' => 10 // Fallback total timeout for older PHP versions ] ]); // Pass the context to file_get_html $html = file_get_html('https://your-target-site.com', false, $context); // Immediately handle failed requests if (!$html) { error_log("Failed to fetch page: Network timeout or unreachable server"); // Exit early instead of letting the script hang return; // Or die(), depending on your script structure }
This will cut off the request as soon as the timeout thresholds are hit, preventing the script from hitting the 30-second max execution limit.
2. Switch to CURL for More Control
If you want even better reliability (and support for things like redirects, SSL, and detailed error logging), replace file_get_html with a CURL-based fetch, then pass the raw content to str_get_html (simple_html_dom’s string-parsing counterpart).
Here’s a reusable CURL function:
function scrape_with_curl($url, $timeout = 10) { $ch = curl_init($url); // Configure CURL options curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, $timeout); curl_setopt($ch, CURLOPT_TIMEOUT, $timeout); curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true); // Handle 3xx redirects curl_setopt($ch, CURLOPT_SSL_VERIFYPEER, false); // Optional: Disable SSL check (use cautiously) $page_content = curl_exec($ch); $curl_error = curl_error($ch); curl_close($ch); // Log errors for debugging if ($curl_error) { error_log("CURL Scrape Error: " . $curl_error); return false; } return $page_content; } // Usage example $page_content = scrape_with_curl('https://your-target-site.com', 10); if ($page_content) { $html = str_get_html($page_content); // Your parsing logic here } else { echo "Failed to retrieve page content"; }
CURL’s timeout controls are more robust than the default stream context, making this my go-to method for production scraping.
3. Adjust Script Execution Time (If Necessary)
If you’re scraping multiple pages and need a bit more breathing room, you can extend the script’s max execution time—but only do this alongside the timeout settings above (don’t rely on it alone):
// Set max execution time to 60 seconds (0 = unlimited, not recommended) set_time_limit(60);
This prevents the script from dying mid-scrape if you have a queue of pages, but the timeout controls from steps 1 or 2 will still stop individual slow requests from dragging things down.
4. Fix Your Error Handling
Your earlier attempt to detect warnings might not have worked because the script was still waiting for the network request to finish. By using the timeout context or CURL, the request will fail fast, and you can exit the function immediately instead of letting it run until the script times out.
Avoid relying on warning detection alone—proactively cut off slow requests with the timeout settings above.
内容的提问来源于stack exchange,提问作者Andy Sanchez

