You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在index.html中使用Node.js脚本?附Reddit爬虫scraper.js代码

如何在index.html中使用Node.js的Reddit抓取脚本

Hey,你得先搞清楚一个核心点:Node.js脚本是运行在服务器/本地Node环境里的,没法直接在浏览器的index.html中直接执行。原因有俩:一是浏览器没有Node.js的核心模块(比如fs、request这些),二是浏览器的跨域安全限制会阻止你直接请求Reddit的页面。

不过别担心,有几种靠谱的方法能让你在前端页面里用上这个抓取逻辑,下面给你一步步讲最常用的方案:

方案一:搭建后端API(网页场景首选)

我们可以把抓取脚本改成一个后端接口,然后前端页面通过AJAX请求这个接口获取数据。

步骤1:准备依赖

首先确保你已经安装了request、cheerio,再安装Express框架来做后端服务:

npm install express request cheerio

步骤2:修改scraper.js为后端接口

把你的抓取逻辑封装成一个HTTP接口,同时让后端托管前端的index.html:

const express = require('express');
const request = require('request');
const cheerio = require('cheerio');
const app = express();
const port = 3000;

// 定义抓取Reddit的接口
app.get('/scrape-reddit', (req, res) => {
  request("https://www.reddit.com/r/all", function(error, response, body) {
    if(error) {
      console.log("Error: " + error);
      return res.status(500).send('抓取数据时出错了');
    }
    // 检查Reddit的响应状态
    if(response.statusCode !== 200) {
      return res.status(400).send('请求Reddit失败,请稍后重试');
    }
    const $ = cheerio.load(body);
    const posts = [];
    // 提取帖子数据(补全你之前没写完的逻辑)
    $('div#siteTable > div.link').each(function( index ) {
      const title = $(this).find('p.title > a.title').text().trim();
      const postUrl = $(this).find('p.title > a.title').attr('href');
      // 只收集有标题的帖子
      if(title) {
        posts.push({ title, url: postUrl });
      }
    });
    // 把数据以JSON格式返回给前端
    res.json(posts);
  });
});

// 托管前端静态文件(把index.html放在public文件夹里)
app.use(express.static('public'));

// 启动服务器
app.listen(port, () => {
  console.log(`服务器已启动,访问 http://localhost:${port} 即可查看页面`);
});

步骤3:创建前端index.html

在项目根目录新建public文件夹,里面创建index.html,用fetch请求后端接口:

<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <title>Reddit热门帖子</title>
  <style>
    ul { list-style: none; padding: 0; }
    li { margin: 10px 0; padding: 8px; border-bottom: 1px solid #eee; }
    a { text-decoration: none; color: #1a73e8; }
    a:hover { text-decoration: underline; }
  </style>
</head>
<body>
  <div class="container">
    <h1>Reddit全站热门帖子</h1>
    <ul id="posts-container"></ul>
  </div>

  <script>
    // 请求后端接口获取抓取的数据
    fetch('/scrape-reddit')
      .then(response => {
        if(!response.ok) throw new Error('请求失败');
        return response.json();
      })
      .then(posts => {
        const container = document.getElementById('posts-container');
        // 把数据渲染到页面上
        posts.forEach(post => {
          const li = document.createElement('li');
          const link = document.createElement('a');
          link.href = post.url;
          link.textContent = post.title;
          link.target = '_blank'; // 新窗口打开链接
          li.appendChild(link);
          container.appendChild(li);
        });
      })
      .catch(err => {
        console.error('加载数据出错:', err);
        document.getElementById('posts-container').textContent = '加载失败,请检查服务器是否运行';
      });
  </script>
</body>
</html>

步骤4:运行项目

在终端执行:

node scraper.js

然后打开浏览器访问http://localhost:3000,就能看到Reddit的热门帖子列表了。

方案二:用Electron做桌面应用(桌面场景)

如果你想做一个桌面应用,而不是网页,可以用Electron——它允许你在桌面应用中同时使用Node.js和前端页面。你可以通过preload脚本把Node.js的抓取逻辑暴露给前端页面,不过这个方案更适合桌面端,而非网页场景。


内容的提问来源于stack exchange,提问作者Dan Hessler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:21:59