You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Express.js中实现垃圾邮件过滤器,基于键值关键词筛选对象

简易垃圾邮件过滤器Express集成实现方案

1 基础环境准备

先初始化Express项目:

  • 执行项目初始化命令:npm init -y
  • 安装Express依赖:npm install express

2 关键词规则过滤实现(满足你当前的筛选需求)

你要求的id===1且message包含指定三个关键词的筛选逻辑可以直接通过Express接口实现:

const express = require('express')
const app = express()
app.use(express.json())

// 定义垃圾邮件关键词组合
const SPAM_KEYWORDS = ['lottery', 'win', 'millionaire']

// 邮件过滤接口,收到邮件后调用该接口传入邮件数据即可
app.get('/filter-email', (req, res) => {
  // 从请求参数中获取邮件列表
  const emailList = JSON.parse(req.query.emailList)
  // 执行筛选
  const spamEmails = emailList.filter(item => {
    // 先匹配id条件
    if (item.id !== '1') return false
    const messageLower = item.message.toLowerCase()
    // 匹配所有关键词
    return SPAM_KEYWORDS.every(keyword => messageLower.includes(keyword.toLowerCase()))
  })
  // 返回结果:区分垃圾邮件和正常邮件
  res.send({
    spamEmails,
    normalEmails: emailList.filter(item => !spamEmails.includes(item))
  })
})

// 启动服务,端口可自行修改
app.listen(3000, () => console.log('过滤服务已启动,端口3000'))

测试时将你给出的示例邮件数组作为emailList参数传入接口,即可正确筛选出第一条为垃圾邮件。

3 集成朴素贝叶斯算法实现智能过滤

如果需要替代固定关键词规则,实现更灵活的垃圾邮件识别,可以按以下步骤集成朴素贝叶斯:

3.1 实现简易朴素贝叶斯分类器

class NaiveBayes {
  constructor() {
    // 分别统计垃圾邮件(spam)和正常邮件(ham)的词频、样本数、总词数
    this.wordMap = { spam: {}, ham: {} }
    this.sampleCount = { spam: 0, ham: 0 }
    this.totalWordCount = { spam: 0, ham: 0 }
  }
  // 简单分词逻辑,可根据需求优化
  #tokenize(text) {
    return text.toLowerCase().match(/[a-zA-Z]+/g) || []
  }
  // 模型训练方法,传入文本和标注分类(spam/ham)
  train(text, category) {
    const tokens = this.#tokenize(text)
    this.sampleCount[category]++
    tokens.forEach(token => {
      this.wordMap[category][token] = (this.wordMap[category][token] || 0) + 1
      this.totalWordCount[category]++
    })
  }
  // 预测方法,返回文本分类结果
  predict(text) {
    const tokens = this.#tokenize(text)
    const totalSample = this.sampleCount.spam + this.sampleCount.ham
    // 计算先验概率,用对数避免下溢
    let spamProb = Math.log(this.sampleCount.spam / totalSample)
    let hamProb = Math.log(this.sampleCount.ham / totalSample)
    tokens.forEach(token => {
      // 拉普拉斯平滑处理零概率问题
      const spamWordProb = (this.wordMap.spam[token] || 0 + 1) / (this.totalWordCount.spam + 2)
      const hamWordProb = (this.wordMap.ham[token] || 0 + 1) / (this.totalWordCount.ham + 2)
      spamProb += Math.log(spamWordProb)
      hamProb += Math.log(hamWordProb)
    })
    return spamProb > hamProb ? 'spam' : 'ham'
  }
}

3.2 集成到Express服务

在服务启动阶段完成模型训练,接口中直接调用预测方法即可:

// 初始化分类器并训练
const classifier = new NaiveBayes()
// 标注训练样本,样本越多准确率越高,可自行扩展
const trainSamples = [
  { text: 'You have a chance to win a lottery and be a millionaire', category: 'spam' },
  { text: 'Win free iphone now click the link', category: 'spam' },
  { text: 'Claim your 10 million lottery prize today', category: 'spam' },
  { text: 'hello how are you doing', category: 'ham' },
  { text: 'Team meeting will be held at 2pm tomorrow', category: 'ham' },
  { text: 'Please review the attached document before Friday', category: 'ham' }
]
trainSamples.forEach(item => classifier.train(item.text, item.category))

// 基于贝叶斯的过滤接口
app.get('/filter-email-bayes', (req, res) => {
  const emailList = JSON.parse(req.query.emailList)
  const spamEmails = emailList.filter(item => {
    return item.id === '1' && classifier.predict(item.message) === 'spam'
  })
  res.send({ spamEmails, normalEmails: emailList.filter(item => !spamEmails.includes(item)) })
})

后续优化方向

  • 如果需要自动接收邮件,可以后续集成POP3/IMAP工具包主动拉取邮件,无需手动调用接口传入数据
  • 朴素贝叶斯的准确率和训练样本量正相关,后续可以持续补充标注样本提升识别效果

内容的提问来源于stack exchange,提问作者Brute

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 18:12:02