You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Express.js的SPA SEO:如何向搜索引擎提交页面

基于Express实现SPA的SEO优化方案

针对你的需求,我们可以通过检测请求的User-Agent区分爬虫与普通用户:对爬虫返回填充好元标签的静态HTML,对普通用户返回原SPA入口文件,具体实现步骤如下:

1. 安装必要依赖

需要用cheerio解析修改HTML模板中的元标签,用axios调用你的Rest API获取元数据(也可替换为你项目中已有的HTTP请求库):

npm install cheerio axios

2. 修改Express服务代码

基于你提供的示例,改造后的完整实现代码如下:

const express = require('express');
const path = require('path');
const fs = require('fs');
const cheerio = require('cheerio');
const axios = require('axios');

const app = express();
const port = process.env.PORT || 8080;
// 提前读取index.html模板并缓存,避免重复文件IO
const indexTemplate = fs.readFileSync(path.join(__dirname, 'public/index.html'), 'utf8');

// 常见爬虫User-Agent关键词列表,可按需补充
const crawlerAgents = [
  'googlebot', 'bingbot', 'slurp', 'duckduckbot', 'baiduspider',
  'yandexbot', 'sogou', 'exabot', 'facebot', 'ia_archiver'
];

// 判断请求是否来自爬虫的工具函数
function isCrawler(req) {
  const userAgent = req.headers['user-agent']?.toLowerCase() || '';
  return crawlerAgents.some(agent => userAgent.includes(agent));
}

// 匹配所有GET请求的通用路由
app.get('*', async (req, res) => {
  // 普通用户请求:直接返回原SPA入口文件
  if (!isCrawler(req)) {
    return res.sendFile(path.join(__dirname, 'public/index.html'));
  }

  // 爬虫请求:调用Rest API获取当前页面元数据并填充模板
  try {
    // 替换为你的Rest API地址,根据请求路径拉取对应页面的元数据
    const apiResponse = await axios.get(`https://your-rest-api-domain.com/api/meta${req.path}`);
    const metaData = apiResponse.data; // 假设返回字段包含title、description、ogTitle等

    // 解析HTML模板并替换元标签
    const $ = cheerio.load(indexTemplate);
    $('title').text(metaData.title || '默认标题');
    $('meta[name="description"]').attr('content', metaData.description || '默认描述');
    // 可选:替换Open Graph标签用于社交平台分享
    $('meta[property="og:title"]').attr('content', metaData.ogTitle || metaData.title || '默认标题');
    $('meta[property="og:description"]').attr('content', metaData.ogDescription || metaData.description || '默认描述');
    $('meta[property="og:url"]').attr('content', `https://your-domain.com${req.path}`);

    // 返回填充好元数据的HTML给爬虫
    res.send($.html());
  } catch (error) {
    // API请求失败时,返回带默认元标签的模板,避免影响爬虫抓取
    const $ = cheerio.load(indexTemplate);
    $('title').text('默认标题');
    $('meta[name="description"]').attr('content', '默认描述');
    res.send($.html());
    console.error('获取元数据失败:', error);
  }
});

app.listen(port, () => {
  console.log(`Server running on port ${port}!`);
});

3. 关键细节说明

  • User-Agent检测:通过匹配常见爬虫UA关键词区分请求来源,可根据业务需求补充更多爬虫标识;
  • 模板缓存:提前读取并缓存index.html,减少服务器文件读取开销;
  • 容错处理:API请求失败时返回带默认元标签的页面,确保爬虫能正常抓取内容;
  • 元标签适配:根据你实际index.html中的元标签结构,调整cheerio选择器保证替换准确;
  • 性能优化:可引入lru-cache等工具缓存已获取的元数据,避免重复调用Rest API。

内容的提问来源于stack exchange,提问作者Александр

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 15:13:16