如何限制页面访问,仅允许社交爬虫与Google爬虫获取<head>内容生成预览
实现方案
核心思路
通过User-Agent识别爬虫,在PHP认证逻辑中加入判断:对合规爬虫仅返回页面<head>元数据,对普通未认证用户则拦截并要求认证,不暴露页面主体内容。
具体步骤
1. 构建爬虫User-Agent白名单
收集主流社交媒体爬虫和Google爬虫的UA标识,常见的包括:
- Googlebot(Google搜索爬虫)
- FacebookExternalHit(Facebook链接预览爬虫)
- Twitterbot(Twitter/X链接预览爬虫)
- LinkedInBot(LinkedIn链接预览爬虫)
- Slackbot(Slack链接预览爬虫)
2. 在PHP认证网关中加入爬虫判断逻辑
在页面最顶部(输出任何内容之前)加入以下逻辑:
<?php // 定义爬虫UA白名单 $allowedCrawlers = [ 'Googlebot', 'FacebookExternalHit', 'Twitterbot', 'LinkedInBot', 'Slackbot' ]; // 获取当前请求的User-Agent $userAgent = $_SERVER['HTTP_USER_AGENT'] ?? ''; // 检查是否为允许的爬虫 $isAllowedCrawler = false; foreach ($allowedCrawlers as $crawler) { if (stripos($userAgent, $crawler) !== false) { $isAllowedCrawler = true; break; } } // 处理逻辑 if (!$isAllowedCrawler) { // 非爬虫用户,检查认证状态 session_start(); if (!isset($_SESSION['authenticated']) || !$_SESSION['authenticated']) { // 未认证,跳转或输出认证提示 header('Location: /login.php'); exit; } } // 后续页面渲染逻辑 ?> <!DOCTYPE html> <html> <head> <title>页面标题</title> <meta name="description" content="页面描述"> <meta property="og:title" content="OG标题"> <meta property="og:description" content="OG描述"> <meta property="og:image" content="/path/to/image.jpg"> <!-- 其他元数据 --> </head> <body> <?php if (!$isAllowedCrawler): ?> <!-- 已认证用户看到的页面主体内容 --> <div class="content"> <!-- 页面内容 --> </div> <?php endif; ?> </body> </html>
3. 增强安全性(可选)
- 结合爬虫官方公布的IP范围进行验证(比如Google公开的Googlebot IP段),降低伪造User-Agent绕过判断的风险。
- 若元数据为动态生成,确保爬虫请求时能正确拉取对应页面的标题、描述等信息。
注意事项
- 不要仅依赖
$_SERVER['HTTP_USER_AGENT']判断,UA可被伪造,但主流爬虫的标识相对稳定,结合IP验证能提升安全性。 - 必须在输出任何HTML内容前完成认证和爬虫判断,避免触发"headers already sent"错误。
内容的提问来源于stack exchange,提问作者Arshavin
相关产品推荐
相关产品推荐

