You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

curl_multi_exec执行过慢求助:45k批量URL请求优化问题

Optimizing curl_multi_exec for 45k URL Requests

Hey there! Let's tackle this slow response time issue you're facing when processing 45k URLs in 200-item batches with curl_multi_exec. I’ve worked on similar large-scale HTTP request tasks before, so here are some practical, actionable optimizations to speed things up:

1. Reuse CURL Handles Instead of Creating New Ones

Creating 200 new CURL handles for every batch is a huge resource drain. Instead, initialize a pool of persistent handles once, then just update the target URL and necessary options for each batch. This cuts down on the overhead of repeated curl_init() and curl_close() calls.

Example snippet for handle reuse:

// Initialize a handle pool once at the start
$handlePool = [];
for ($i = 0; $i < 200; $i++) {
    $ch = curl_init();
    // Set persistent options here (keep-alive, timeouts, etc.)
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_TCP_KEEPALIVE, 1);
    curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 10);
    $handlePool[] = $ch;
}

// For each batch of URLs
foreach ($urlBatches as $batch) {
    $mh = curl_multi_init();
    foreach ($batch as $index => $url) {
        $ch = $handlePool[$index];
        curl_setopt($ch, CURLOPT_URL, $url);
        curl_multi_add_handle($mh, $ch);
    }

    // Execute the batch (with optimized select logic)
    do {
        $status = curl_multi_exec($mh, $active);
        $selectStatus = curl_multi_select($mh);
        if ($selectStatus === -1) {
            usleep(100); // Prevent busy waiting
        }
    } while ($active > 0 && $status === CURLM_OK);

    // Process responses and remove handles
    foreach ($batch as $index => $url) {
        $ch = $handlePool[$index];
        $response = curl_multi_getcontent($ch);
        // Process $response here (write to file/db instead of storing in array!)
        curl_multi_remove_handle($mh, $ch);
    }
    curl_multi_close($mh);
}

// Cleanup handles at the end
foreach ($handlePool as $ch) {
    curl_close($ch);
}

2. Tune Core CURL Configuration

Tweak these options to reduce overhead and speed up connections:

  • Enable HTTP/2 or Keep-Alive: If the target server supports HTTP/2, use curl_setopt($ch, CURLOPT_HTTP_VERSION, CURL_HTTP_VERSION_2_0); to multiplex requests over a single TCP connection. For HTTP/1.1, ensure CURLOPT_TCP_KEEPALIVE is enabled to reuse connections.
  • Set Strict Timeouts: Avoid hanging on slow requests with CURLOPT_TIMEOUT (total request timeout) and CURLOPT_CONNECTTIMEOUT (connection timeout). Start with 10-30 seconds depending on your use case.
  • Disable Unnecessary Features: Turn off CURLOPT_VERBOSE, CURLOPT_FOLLOWLOCATION (if you don’t need redirects), and only disable CURLOPT_SSL_VERIFYPEER if you fully trust the target servers (not recommended for production).

3. Adjust Batch Size

200 might not be the sweet spot—system limits (like open file descriptors) or network bandwidth could be throttling your requests. Test smaller batches (100, 150) or slightly larger ones (250) to find the optimal size. You can check your system’s file descriptor limit with ulimit -n on Linux/macOS.

4. Optimize the Execution Loop

A common mistake is using a fixed sleep() in the curl_multi_exec loop, which wastes time. Use curl_multi_select() properly to wait for activity without busy waiting:

do {
    $status = curl_multi_exec($mh, $active);
    $selectResult = curl_multi_select($mh);
    // If select returns -1, sleep briefly to avoid 100% CPU usage
    if ($selectResult === -1) {
        usleep(100);
    }
} while ($active > 0 && $status === CURLM_OK);

5. Reduce Memory Overhead

Storing all 45k responses in an array will bloat your memory and trigger frequent garbage collection. Instead, process each response immediately as you retrieve it—write it to a file, insert it into a database, or stream it to another service. This keeps memory usage low and prevents slowdowns from memory pressure.

6. Optimize DNS Resolution

Enable DNS caching to avoid repeated lookups for the same domain:

curl_setopt($ch, CURLOPT_DNS_CACHE_TIMEOUT, 3600); // Cache for 1 hour

If you’re hitting many different domains, consider using a local DNS resolver (like dnsmasq) to speed up lookups.

Try implementing these changes one at a time and measure the performance after each tweak—you should see a significant reduction in total execution time. If you run into specific issues with any of these steps, feel free to share more details!

内容的提问来源于stack exchange,提问作者Iqra Raees

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:12:22