curl_setopt问题:跳过不可访问域名,解决连接重置报错
问题
当从DOMAINS.json读取多个域名处理时,只要有一个域名无法访问(触发ERR_TIMED_OUT或ERR_CONNECTION_RESET),脚本就会直接崩溃终止。需要修改代码,让脚本自动跳过这类不可访问的域名,继续处理剩余条目。
现有核心代码片段:
$sites = file_get_contents('./DOMAINS.json'); $sites = json_decode($sites); function get_html_title($html){ preg_match("/\<title.*\>(.*)\<\/title\>/isU", $html, $matches); return $matches[1]; } function get_redirect_target($url) { $ch = curl_init($url); curl_setopt($ch, CURLOPT_HEADER, 1); curl_setopt($ch, CURLOPT_NOBODY, 1); curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 1); // The maximum number of seconds to allow cURL functions to execute. curl_setopt($ch, CURLOPT_TIMEOUT, 8); //timeout in seconds $headers = curl_exec($ch); curl_close($ch); }
解决办法
要实现跳过异常域名,需在三个核心环节添加容错处理:
1. 给cURL请求增加错误捕获
在get_redirect_target函数中加入错误检测逻辑,请求失败时返回明确标记,避免脚本抛出致命错误:
function get_redirect_target($url) { $ch = curl_init($url); curl_setopt($ch, CURLOPT_HEADER, 1); curl_setopt($ch, CURLOPT_NOBODY, 1); curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 1); curl_setopt($ch, CURLOPT_TIMEOUT, 8); curl_setopt($ch, CURLOPT_FOLLOWLOCATION, false); curl_setopt($ch, CURLOPT_SSL_VERIFYPEER, false); // 按需关闭SSL验证,避免证书问题导致失败 curl_setopt($ch, CURLOPT_SSL_VERIFYHOST, false); $headers = curl_exec($ch); // 检查cURL请求是否出错 if (curl_errno($ch)) { curl_close($ch); return false; // 请求失败返回false } curl_close($ch); // 解析跳转目标(原有逻辑保留) if (preg_match('/Location: (.*)\n/', $headers, $matches)) { return trim($matches[1]); } return $url; }
2. 遍历域名时跳过异常条目
在处理每个域名的循环中,判断请求结果,若失败则直接跳过当前域名:
$sites = file_get_contents('./DOMAINS.json'); $sites = json_decode($sites, true); // 转为数组,处理更灵活 // 遍历所有域名 foreach ($sites as $site) { $target = get_redirect_target($site); if (!$target) { echo "跳过不可访问域名: {$site}\n"; continue; // 跳过当前条目,继续处理下一个 } // 后续处理示例(获取页面标题) $html = file_get_contents($target); if ($html) { $title = get_html_title($html); echo "域名: {$site} | 标题: {$title}\n"; } else { echo "无法获取{$site}的页面内容\n"; } }
3. 修复get_html_title的容错问题
原函数在未匹配到title标签时会因数组索引不存在报错,增加判断逻辑:
function get_html_title($html){ if (preg_match("/\<title.*\>(.*)\<\/title\>/isU", $html, $matches)) { return trim($matches[1]); } return '无标题'; // 匹配失败返回默认值 }
额外提示
- 如果
file_get_contents也经常因网络问题报错,可替换为带错误处理的cURL请求,进一步提升稳定性。 CURLOPT_SSL_VERIFYPEER和CURLOPT_SSL_VERIFYHOST的关闭是为了避免SSL证书问题导致的请求失败,生产环境若需严格验证可开启。
内容的提问来源于stack exchange,提问作者amir salarr
相关产品推荐
相关产品推荐

