如何使用curl解析20+站点数据?解决504网关超时问题
解决WordPress插件Curl批量请求504网关超时问题
批量请求29个站点触发504错误,核心原因是总请求耗时超过了服务器网关或PHP的超时限制,加上当前串行请求的方式效率太低,拖长了整体执行时间。以下是具体解决办法:
1. 优化单Curl请求的参数
给每个请求加上超时限制,避免单个慢响应站点拖垮整个任务,同时减少不必要的开销:
$ch = curl_init(); curl_setopt($ch, CURLOPT_URL, $url); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); // 设置连接和总超时时间,按需调整 curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 10); // 连接超时10秒 curl_setopt($ch, CURLOPT_TIMEOUT, 15); // 整个请求超时15秒 // 若目标站点无需严格SSL验证,可关闭以减少握手耗时 curl_setopt($ch, CURLOPT_SSL_VERIFYPEER, false); curl_setopt($ch, CURLOPT_SSL_VERIFYHOST, false); $html = curl_exec($ch); curl_close($ch);
另外可以复用Curl句柄,不用每次都初始化和关闭,减少连接建立的开销:
// 只初始化一次句柄 $ch = curl_init(); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 10); curl_setopt($ch, CURLOPT_TIMEOUT, 15); foreach ($all_urls as $url) { curl_setopt($ch, CURLOPT_URL, $url); $html = curl_exec($ch); // 处理数据逻辑... } curl_close($ch);
2. 改用并行Curl请求
串行请求的总耗时是每个请求的时间总和,改用并行请求能大幅压缩总耗时,用Curl Multi实现:
$mh = curl_multi_init(); $handles = []; // 先把所有请求添加到多句柄中 foreach ($all_urls as $url) { $ch = curl_init($url); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 10); curl_setopt($ch, CURLOPT_TIMEOUT, 15); curl_multi_add_handle($mh, $ch); $handles[] = $ch; } // 执行并行请求 $running = null; do { curl_multi_exec($mh, $running); curl_multi_select($mh); } while ($running > 0); // 逐个获取结果并清理 foreach ($handles as $ch) { $html = curl_multi_getcontent($ch); // 处理数据逻辑... curl_multi_remove_handle($mh, $ch); curl_close($ch); } curl_multi_close($mh);
3. 调整服务器的超时配置
- PHP层面:如果PHP执行时间不够,修改
max_execution_time,可以在代码开头临时设置:
也可以直接修改ini_set('max_execution_time', 120); // 设为120秒,按需调整php.ini文件中的对应参数。 - Web服务器层面:
- 若用Nginx,修改站点配置里的
proxy_read_timeout,比如设为120秒:location / { proxy_read_timeout 120s; } - 若用Apache,修改
httpd.conf里的Timeout指令:Timeout 120
- 若用Nginx,修改站点配置里的
4. 拆分任务分批执行
把29个站点分成多批(比如每批10个),用WordPress的定时任务(WP Cron)分批跑,避免单次请求耗时过长:
// 注册定时任务 add_action('wp', 'register_crawl_batch_task'); function register_crawl_batch_task() { if (!wp_next_scheduled('crawl_batch_event')) { wp_schedule_event(time(), 'hourly', 'crawl_batch_event'); } } // 执行分批爬取 add_action('crawl_batch_event', 'run_crawl_batch'); function run_crawl_batch() { $all_urls = ['站点1地址', '站点2地址', ..., '站点29地址']; $batch_size = 10; // 记录当前执行到第几批 $current_batch = get_option('crawl_current_batch', 0); $start = $current_batch * $batch_size; $batch_urls = array_slice($all_urls, $start, $batch_size); // 处理当前批次的请求(用优化后的串行或并行方法) foreach ($batch_urls as $url) { $ch = curl_init($url); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 10); curl_setopt($ch, CURLOPT_TIMEOUT, 15); $html = curl_exec($ch); curl_close($ch); // 处理数据逻辑... } // 更新批次索引,完成所有批次后重置 $next_batch = $current_batch + 1; if ($next_batch * $batch_size >= count($all_urls)) { delete_option('crawl_current_batch'); } else { update_option('crawl_current_batch', $next_batch); } }
内容的提问来源于stack exchange,提问作者Shahariar Ashik
相关产品推荐
相关产品推荐

