如何在Vercel部署的Next.js应用中仅禁止子域名爬虫?
如何在Vercel部署的Next.js应用中针对不同域名配置robots.txt
因为静态robots.txt无法识别请求来源的域名,所以必须通过动态生成robots内容的方式实现需求,以下是两种Next.js路由方案的具体操作:
Pages Router 方案(Next.js 12及更早版本)
- 删除项目根目录
public文件夹下的静态robots.txt文件 - 在
pages/api目录下新建robots.ts(或.js)文件,编写如下代码:
import type { NextApiRequest, NextApiResponse } from 'next' export default function handler(req: NextApiRequest, res: NextApiResponse) { const host = req.headers.host // 主域名的爬虫规则:允许所有抓取 const mainRobots = `User-agent: * Allow: / Sitemap: https://www.example.com/sitemap.xml` // 子域名的爬虫规则:完全禁止抓取 const testerRobots = `User-agent: * Disallow: /` // 根据访问域名返回对应规则 if (host === 'testers.example.com') { res.setHeader('Content-Type', 'text/plain') res.status(200).send(testerRobots) } else { res.setHeader('Content-Type', 'text/plain') res.status(200).send(mainRobots) } }
部署后,访问https://testers.example.com/robots.txt会返回禁止抓取的规则,主域名则返回允许规则。
App Router 方案(Next.js 13+)
- 删除
public文件夹下的静态robots.txt文件 - 在
app目录下新建robots.txt文件夹,里面创建route.ts文件,编写如下代码:
import { NextRequest, NextResponse } from 'next/server' export function GET(request: NextRequest) { const host = request.headers.get('host') const mainRobots = `User-agent: * Allow: / Sitemap: https://www.example.com/sitemap.xml` const testerRobots = `User-agent: * Disallow: /` let robotsContent = mainRobots if (host === 'testers.example.com') { robotsContent = testerRobots } return new NextResponse(robotsContent, { headers: { 'Content-Type': 'text/plain', }, }) }
部署完成后,两个域名的robots.txt会根据访问来源返回对应规则,爬虫会遵守各自域名的限制。
注意事项
- 确保
host判断准确,生产环境下Vercel返回的host就是纯域名(不带端口),无需额外处理 - 如果需要更精细的规则(比如只禁止特定爬虫),可以修改
User-agent和对应的Allow/Disallow内容
内容的提问来源于stack exchange,提问作者Anand Yadav
相关产品推荐
相关产品推荐

