You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

simple_html_dom服务器端爬取forebet网站失败,本地运行正常求助

PHP爬虫部署服务器后爬取forebet.com报403 Forbidden的解决方法

我开发了一个PHP爬虫,用于爬取https://www.forebet.com/en/values的页面标题,代码如下:

require 'simple_html_dom.php';
$html=file_get_html("https://www.forebet.com/en/values");
echo $html->find('title', 0)->plaintext;

本地运行正常,但部署到服务器后出现HTTP 500错误。开启错误显示后,发现两个问题:

  • 警告:file_get_contents(https://www.forebet.com/en/values): failed to open stream: HTTP request failed! HTTP/1.1 403 Forbidden,出现在simple_html_dom.php的84行(对应代码$contents = file_get_contents( $url, $use_include_path, $context, $offset, $maxLen);)
  • 致命错误:Uncaught Error: Call to a member function find() on bool,出现在php.php的16行(对应代码echo $html->find('title', 0)->plaintext;)

测试爬取google.com时服务器端运行正常。


问题原因

服务器发起的请求被forebet.com的反爬机制拦截了——默认file_get_contents的请求头会暴露是PHP脚本发起的请求,而本地环境的请求头通常带有浏览器标识,因此未被拦截。

解决方案

1. 给file_get_contents添加浏览器请求头

通过创建请求上下文,模拟浏览器的User-Agent标识,绕过反爬检测:

require 'simple_html_dom.php';

// 构造模拟浏览器的请求头
$context = stream_context_create([
    'http' => [
        'header' => "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36\r\n"
    ]
]);

// 先获取页面内容,再传入simple_html_dom解析
$htmlContent = file_get_contents("https://www.forebet.com/en/values", false, $context);
if ($htmlContent) {
    $html = str_get_html($htmlContent);
    if ($html && $title = $html->find('title', 0)) {
        echo $title->plaintext;
    } else {
        echo "无法获取标题";
    }
} else {
    echo "请求页面失败";
}

2. 改用cURL发起请求(更稳定)

cURL相比file_get_contents更灵活,能更好适配反爬策略,示例代码:

require 'simple_html_dom.php';

$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, "https://www.forebet.com/en/values");
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
// 设置浏览器User-Agent
curl_setopt($ch, CURLOPT_USERAGENT, "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36");
// 可选:添加语言请求头,更贴近真实浏览器
curl_setopt($ch, CURLOPT_HTTPHEADER, [
    "Accept-Language: en-US,en;q=0.9"
]);

$htmlContent = curl_exec($ch);
$httpCode = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

if ($httpCode == 200 && $htmlContent) {
    $html = str_get_html($htmlContent);
    if ($html && $title = $html->find('title', 0)) {
        echo $title->plaintext;
    } else {
        echo "无法解析标题";
    }
} else {
    echo "请求失败,状态码:$httpCode";
}

3. 强制添加错误处理逻辑

无论用哪种请求方式,都要先判断请求结果和解析对象是否有效,避免致命错误:

require 'simple_html_dom.php';
$html=file_get_html("https://www.forebet.com/en/values");

// 先判断页面是否获取成功
if ($html !== false) {
    $title = $html->find('title', 0);
    // 再判断标题是否存在
    if ($title) {
        echo $title->plaintext;
    } else {
        echo "未找到标题";
    }
} else {
    echo "无法获取页面内容";
}

内容的提问来源于stack exchange,提问作者Muhammad Umer Farooq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 10:50:24