Spring Boot MongoDB如何反向检测_id是否为输入字符串的子串?
解决方案:敏感词检测与替换
核心问题分析
你之前的查询逻辑完全搞反了:原查询是判断_id是否包含输入字符串,而实际需求是输入字符串是否包含_id对应的敏感词。下面提供两种可行方案,按需选择。
方案一:Java端批量处理(适合敏感词数量较少场景)
先从MongoDB拉取所有敏感词,再在Java层遍历检查输入字符串是否包含敏感词,最后完成替换。
步骤1:查询所有敏感词
通过MongoOperations或MongoRepository获取全量敏感词:
// 方式1:用MongoOperations List<BadWord> badWords = mongoOperations.findAll(BadWord.class); Set<String> sensitiveWords = badWords.stream() .map(BadWord::getWord) .collect(Collectors.toSet()); // 方式2:用MongoRepository(需先定义接口) // public interface BadWordRepository extends MongoRepository<BadWord, String> {} List<BadWord> badWords = badWordRepository.findAll();
步骤2:敏感词替换
遍历敏感词,将输入字符串中匹配的部分替换为等长的*(支持大小写不敏感):
public String filterSensitiveWords(String input, Set<String> sensitiveWords) { if (input == null || sensitiveWords.isEmpty()) { return input; } // 构造正则表达式:转义敏感词特殊字符,避免正则语法冲突 String regex = String.join("|", sensitiveWords.stream() .map(Pattern::quote) .collect(Collectors.toList())); Pattern pattern = Pattern.compile(regex, Pattern.CASE_INSENSITIVE); Matcher matcher = pattern.matcher(input); // 替换为与原敏感词等长的* return matcher.replaceAll(matchResult -> "*".repeat(matchResult.group().length()) ); }
测试效果
- 输入
badword1→ 输出********* - 输入
badword1badword2→ 输出****************** - 输入
badword123→ 输出*********
方案二:MongoDB端匹配(适合敏感词数量较多场景)
利用MongoDB的$regexMatch聚合操作,直接在数据库层筛选出输入字符串包含的敏感词,减少Java端内存占用。
步骤1:构造反向匹配查询
使用Criteria结合$expr实现输入字符串对_id的包含判断:
public List<BadWord> findMatchingSensitiveWords(String input) { Query query = new Query(); // 用$regexMatch判断输入字符串是否包含当前文档的_id query.addCriteria(Criteria.where("$expr").is( new Document("$regexMatch", new Document("input", input) .append("regex", "$_id") .append("options", "i") // 忽略大小写,不需要则删除该参数 ) )); return mongoOperations.find(query, BadWord.class); }
步骤2:执行替换
拿到数据库返回的匹配敏感词后,复用方案一中的替换逻辑即可。
注意事项
- 性能优化:敏感词数量过万时,优先选方案二;数量较少时方案一更简单。
- 特殊字符处理:必须用
Pattern.quote()转义敏感词中的正则特殊字符(如+、*、.),避免匹配逻辑出错。 - 大小写控制:根据业务需求调整
options参数,不需要忽略大小写则去掉"i"。
内容的提问来源于stack exchange,提问作者VK321
相关产品推荐
相关产品推荐

