You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Web Scraping检测New World官网最新发布的新闻文章

解决方法

核心逻辑是通过存储历史爬取到的文章链接,每次新爬取后做对比筛选出新增条目,修改后的代码如下:

const axios = require('axios')
const cheerio = require('cheerio')
const express = require('express')
const PORT = 3000
const URL = 'https://www.newworld.com/en-us/news'
const splitURL = URL.slice(0, 24)
// 新增:存储已爬取过的文章链接,用Set实现O(1)复杂度的存在性查询
const existedLinks = new Set()

const app = express()

// 注意:setInterval第二个参数是毫秒,原代码写PORT=3000即3秒爬一次,频率太高容易被反爬拦截,可自行调整,比如10分钟设为600000
setInterval(() => {
    axios(URL)
    .then(response => {
        const html = response.data
        const $ = cheerio.load(html)
        const articles = []
        $('.ags-SlotModule', html).each(function(){
            const link = splitURL + $(this).find('a').attr('href')
            articles.push({
                link
            })
        })
        // 新增:对比筛选新增文章
        const newArticles = articles.filter(item => !existedLinks.has(item.link))
        // 有新增内容时打印
        if (newArticles.length > 0) {
            console.log('检测到新增新闻:')
            newArticles.forEach(article => console.log(article))
            // 将新增链接加入历史集合
            newArticles.forEach(article => existedLinks.add(article.link))
        }
    }).catch(err => console.log(err))
}, 600000) // 这里默认改为10分钟爬一次,可按需调整

app.listen(PORT, () => console.log('server running'))

注意事项

  • 如果需要服务重启后仍能识别之前的历史文章,可以将existedLinks的内容持久化存储到本地JSON文件、Redis或者轻量数据库中,启动时读取数据初始化即可
  • 可以额外爬取文章标题、发布时间等信息,和链接组合判重,避免极端情况下链接重复导致的判断错误

内容的提问来源于stack exchange,提问作者Idiot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 23:36:03