如何在Express.js中实现垃圾邮件过滤器,基于键值关键词筛选对象
简易垃圾邮件过滤器Express集成实现方案
1 基础环境准备
先初始化Express项目:
- 执行项目初始化命令:
npm init -y - 安装Express依赖:
npm install express
2 关键词规则过滤实现(满足你当前的筛选需求)
你要求的id===1且message包含指定三个关键词的筛选逻辑可以直接通过Express接口实现:
const express = require('express') const app = express() app.use(express.json()) // 定义垃圾邮件关键词组合 const SPAM_KEYWORDS = ['lottery', 'win', 'millionaire'] // 邮件过滤接口,收到邮件后调用该接口传入邮件数据即可 app.get('/filter-email', (req, res) => { // 从请求参数中获取邮件列表 const emailList = JSON.parse(req.query.emailList) // 执行筛选 const spamEmails = emailList.filter(item => { // 先匹配id条件 if (item.id !== '1') return false const messageLower = item.message.toLowerCase() // 匹配所有关键词 return SPAM_KEYWORDS.every(keyword => messageLower.includes(keyword.toLowerCase())) }) // 返回结果:区分垃圾邮件和正常邮件 res.send({ spamEmails, normalEmails: emailList.filter(item => !spamEmails.includes(item)) }) }) // 启动服务,端口可自行修改 app.listen(3000, () => console.log('过滤服务已启动,端口3000'))
测试时将你给出的示例邮件数组作为emailList参数传入接口,即可正确筛选出第一条为垃圾邮件。
3 集成朴素贝叶斯算法实现智能过滤
如果需要替代固定关键词规则,实现更灵活的垃圾邮件识别,可以按以下步骤集成朴素贝叶斯:
3.1 实现简易朴素贝叶斯分类器
class NaiveBayes { constructor() { // 分别统计垃圾邮件(spam)和正常邮件(ham)的词频、样本数、总词数 this.wordMap = { spam: {}, ham: {} } this.sampleCount = { spam: 0, ham: 0 } this.totalWordCount = { spam: 0, ham: 0 } } // 简单分词逻辑,可根据需求优化 #tokenize(text) { return text.toLowerCase().match(/[a-zA-Z]+/g) || [] } // 模型训练方法,传入文本和标注分类(spam/ham) train(text, category) { const tokens = this.#tokenize(text) this.sampleCount[category]++ tokens.forEach(token => { this.wordMap[category][token] = (this.wordMap[category][token] || 0) + 1 this.totalWordCount[category]++ }) } // 预测方法,返回文本分类结果 predict(text) { const tokens = this.#tokenize(text) const totalSample = this.sampleCount.spam + this.sampleCount.ham // 计算先验概率,用对数避免下溢 let spamProb = Math.log(this.sampleCount.spam / totalSample) let hamProb = Math.log(this.sampleCount.ham / totalSample) tokens.forEach(token => { // 拉普拉斯平滑处理零概率问题 const spamWordProb = (this.wordMap.spam[token] || 0 + 1) / (this.totalWordCount.spam + 2) const hamWordProb = (this.wordMap.ham[token] || 0 + 1) / (this.totalWordCount.ham + 2) spamProb += Math.log(spamWordProb) hamProb += Math.log(hamWordProb) }) return spamProb > hamProb ? 'spam' : 'ham' } }
3.2 集成到Express服务
在服务启动阶段完成模型训练,接口中直接调用预测方法即可:
// 初始化分类器并训练 const classifier = new NaiveBayes() // 标注训练样本,样本越多准确率越高,可自行扩展 const trainSamples = [ { text: 'You have a chance to win a lottery and be a millionaire', category: 'spam' }, { text: 'Win free iphone now click the link', category: 'spam' }, { text: 'Claim your 10 million lottery prize today', category: 'spam' }, { text: 'hello how are you doing', category: 'ham' }, { text: 'Team meeting will be held at 2pm tomorrow', category: 'ham' }, { text: 'Please review the attached document before Friday', category: 'ham' } ] trainSamples.forEach(item => classifier.train(item.text, item.category)) // 基于贝叶斯的过滤接口 app.get('/filter-email-bayes', (req, res) => { const emailList = JSON.parse(req.query.emailList) const spamEmails = emailList.filter(item => { return item.id === '1' && classifier.predict(item.message) === 'spam' }) res.send({ spamEmails, normalEmails: emailList.filter(item => !spamEmails.includes(item)) }) })
后续优化方向
- 如果需要自动接收邮件,可以后续集成POP3/IMAP工具包主动拉取邮件,无需手动调用接口传入数据
- 朴素贝叶斯的准确率和训练样本量正相关,后续可以持续补充标注样本提升识别效果
内容的提问来源于stack exchange,提问作者Brute
相关产品推荐
相关产品推荐

