curl_multi_exec执行过慢求助:45k批量URL请求优化问题
Hey there! Let's tackle this slow response time issue you're facing when processing 45k URLs in 200-item batches with curl_multi_exec. I’ve worked on similar large-scale HTTP request tasks before, so here are some practical, actionable optimizations to speed things up:
1. Reuse CURL Handles Instead of Creating New Ones
Creating 200 new CURL handles for every batch is a huge resource drain. Instead, initialize a pool of persistent handles once, then just update the target URL and necessary options for each batch. This cuts down on the overhead of repeated curl_init() and curl_close() calls.
Example snippet for handle reuse:
// Initialize a handle pool once at the start $handlePool = []; for ($i = 0; $i < 200; $i++) { $ch = curl_init(); // Set persistent options here (keep-alive, timeouts, etc.) curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_TCP_KEEPALIVE, 1); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 10); $handlePool[] = $ch; } // For each batch of URLs foreach ($urlBatches as $batch) { $mh = curl_multi_init(); foreach ($batch as $index => $url) { $ch = $handlePool[$index]; curl_setopt($ch, CURLOPT_URL, $url); curl_multi_add_handle($mh, $ch); } // Execute the batch (with optimized select logic) do { $status = curl_multi_exec($mh, $active); $selectStatus = curl_multi_select($mh); if ($selectStatus === -1) { usleep(100); // Prevent busy waiting } } while ($active > 0 && $status === CURLM_OK); // Process responses and remove handles foreach ($batch as $index => $url) { $ch = $handlePool[$index]; $response = curl_multi_getcontent($ch); // Process $response here (write to file/db instead of storing in array!) curl_multi_remove_handle($mh, $ch); } curl_multi_close($mh); } // Cleanup handles at the end foreach ($handlePool as $ch) { curl_close($ch); }
2. Tune Core CURL Configuration
Tweak these options to reduce overhead and speed up connections:
- Enable HTTP/2 or Keep-Alive: If the target server supports HTTP/2, use
curl_setopt($ch, CURLOPT_HTTP_VERSION, CURL_HTTP_VERSION_2_0);to multiplex requests over a single TCP connection. For HTTP/1.1, ensureCURLOPT_TCP_KEEPALIVEis enabled to reuse connections. - Set Strict Timeouts: Avoid hanging on slow requests with
CURLOPT_TIMEOUT(total request timeout) andCURLOPT_CONNECTTIMEOUT(connection timeout). Start with 10-30 seconds depending on your use case. - Disable Unnecessary Features: Turn off
CURLOPT_VERBOSE,CURLOPT_FOLLOWLOCATION(if you don’t need redirects), and only disableCURLOPT_SSL_VERIFYPEERif you fully trust the target servers (not recommended for production).
3. Adjust Batch Size
200 might not be the sweet spot—system limits (like open file descriptors) or network bandwidth could be throttling your requests. Test smaller batches (100, 150) or slightly larger ones (250) to find the optimal size. You can check your system’s file descriptor limit with ulimit -n on Linux/macOS.
4. Optimize the Execution Loop
A common mistake is using a fixed sleep() in the curl_multi_exec loop, which wastes time. Use curl_multi_select() properly to wait for activity without busy waiting:
do { $status = curl_multi_exec($mh, $active); $selectResult = curl_multi_select($mh); // If select returns -1, sleep briefly to avoid 100% CPU usage if ($selectResult === -1) { usleep(100); } } while ($active > 0 && $status === CURLM_OK);
5. Reduce Memory Overhead
Storing all 45k responses in an array will bloat your memory and trigger frequent garbage collection. Instead, process each response immediately as you retrieve it—write it to a file, insert it into a database, or stream it to another service. This keeps memory usage low and prevents slowdowns from memory pressure.
6. Optimize DNS Resolution
Enable DNS caching to avoid repeated lookups for the same domain:
curl_setopt($ch, CURLOPT_DNS_CACHE_TIMEOUT, 3600); // Cache for 1 hour
If you’re hitting many different domains, consider using a local DNS resolver (like dnsmasq) to speed up lookups.
Try implementing these changes one at a time and measure the performance after each tweak—you should see a significant reduction in total execution time. If you run into specific issues with any of these steps, feel free to share more details!
内容的提问来源于stack exchange,提问作者Iqra Raees

