JS自动审核白名单过滤无法识别大小写变体词汇问题求助
问题根源
现有代码中使用的includes()、split()方法均为大小写敏感匹配逻辑,仅当用户输入内容和白名单词汇的大小写完全一致时才会命中处理规则,因此大小写变体的白名单词汇无法被识别移除。
修复方案
推荐使用正则匹配实现大小写不敏感的全量替换,逻辑更简洁,性能也优于原有filter+reduce+split+join的写法:
let newContent = censorList.whitelist.reduce((content, whiteWord) => { // 转义白名单词汇中的正则特殊字符,避免匹配异常 const escapedWord = whiteWord.replace(/[.*+?^${}()|[\]\\]/g, '\\$&') // 构造带全局匹配、忽略大小写标志的正则 const matchReg = new RegExp(escapedWord, 'gi') return content.replace(matchReg, '') }, message.content)
- 正则的
g标志代表全局匹配,会替换内容中所有命中的白名单词汇,而非仅替换第一个 - 正则的
i标志代表忽略大小写匹配,可命中所有大小写变体的白名单词汇 - 提前对特殊字符做转义处理,可避免白名单词汇包含
.、*、+等正则保留字符时出现逻辑错误
如果你不想使用正则实现,也可以用纯字符串操作的兼容方案:
let newContent = message.content censorList.whitelist.forEach(whiteWord => { const lowerWhiteWord = whiteWord.toLowerCase() let tempLowerContent = newContent.toLowerCase() let matchIndex = tempLowerContent.indexOf(lowerWhiteWord) // 循环匹配所有命中的词汇并移除 while (matchIndex !== -1) { newContent = newContent.slice(0, matchIndex) + newContent.slice(matchIndex + whiteWord.length) tempLowerContent = newContent.toLowerCase() matchIndex = tempLowerContent.indexOf(lowerWhiteWord) } })
内容的提问来源于stack exchange,提问作者Luckie
相关产品推荐
相关产品推荐

