simple_html_dom服务器端爬取forebet网站失败,本地运行正常求助
PHP爬虫部署服务器后爬取forebet.com报403 Forbidden的解决方法
我开发了一个PHP爬虫,用于爬取https://www.forebet.com/en/values的页面标题,代码如下:
require 'simple_html_dom.php'; $html=file_get_html("https://www.forebet.com/en/values"); echo $html->find('title', 0)->plaintext;
本地运行正常,但部署到服务器后出现HTTP 500错误。开启错误显示后,发现两个问题:
- 警告:
file_get_contents(https://www.forebet.com/en/values): failed to open stream: HTTP request failed! HTTP/1.1 403 Forbidden,出现在simple_html_dom.php的84行(对应代码$contents = file_get_contents( $url, $use_include_path, $context, $offset, $maxLen);) - 致命错误:
Uncaught Error: Call to a member function find() on bool,出现在php.php的16行(对应代码echo $html->find('title', 0)->plaintext;)
测试爬取google.com时服务器端运行正常。
问题原因
服务器发起的请求被forebet.com的反爬机制拦截了——默认file_get_contents的请求头会暴露是PHP脚本发起的请求,而本地环境的请求头通常带有浏览器标识,因此未被拦截。
解决方案
1. 给file_get_contents添加浏览器请求头
通过创建请求上下文,模拟浏览器的User-Agent标识,绕过反爬检测:
require 'simple_html_dom.php'; // 构造模拟浏览器的请求头 $context = stream_context_create([ 'http' => [ 'header' => "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36\r\n" ] ]); // 先获取页面内容,再传入simple_html_dom解析 $htmlContent = file_get_contents("https://www.forebet.com/en/values", false, $context); if ($htmlContent) { $html = str_get_html($htmlContent); if ($html && $title = $html->find('title', 0)) { echo $title->plaintext; } else { echo "无法获取标题"; } } else { echo "请求页面失败"; }
2. 改用cURL发起请求(更稳定)
cURL相比file_get_contents更灵活,能更好适配反爬策略,示例代码:
require 'simple_html_dom.php'; $ch = curl_init(); curl_setopt($ch, CURLOPT_URL, "https://www.forebet.com/en/values"); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); // 设置浏览器User-Agent curl_setopt($ch, CURLOPT_USERAGENT, "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"); // 可选:添加语言请求头,更贴近真实浏览器 curl_setopt($ch, CURLOPT_HTTPHEADER, [ "Accept-Language: en-US,en;q=0.9" ]); $htmlContent = curl_exec($ch); $httpCode = curl_getinfo($ch, CURLINFO_HTTP_CODE); curl_close($ch); if ($httpCode == 200 && $htmlContent) { $html = str_get_html($htmlContent); if ($html && $title = $html->find('title', 0)) { echo $title->plaintext; } else { echo "无法解析标题"; } } else { echo "请求失败,状态码:$httpCode"; }
3. 强制添加错误处理逻辑
无论用哪种请求方式,都要先判断请求结果和解析对象是否有效,避免致命错误:
require 'simple_html_dom.php'; $html=file_get_html("https://www.forebet.com/en/values"); // 先判断页面是否获取成功 if ($html !== false) { $title = $html->find('title', 0); // 再判断标题是否存在 if ($title) { echo $title->plaintext; } else { echo "未找到标题"; } } else { echo "无法获取页面内容"; }
内容的提问来源于stack exchange,提问作者Muhammad Umer Farooq
相关产品推荐
相关产品推荐

