MongoDB中检测字符串是否包含指定关键词应使用什么正确函数?
问题原因
$in运算符的使用逻辑完全错误:该运算符用于判断左侧字段的完整值是否等于右侧数组中的某一个元素,不支持判断字符串是否包含子串或匹配正则。- 第一段代码中,你在
$in的右侧数组加入正则也无法生效,$in不支持这种场景下的正则子串匹配,只会判断tweet字段的完整值是否和数组内的某一个正则/字符串完全相等,自然匹配失败。 - 第二段代码中,你直接传入字符串数组,只有当tweet的内容完全是
Tesla、rocket、coffee三者之一时才会命中,包含对应词汇的长文本不会触发匹配。
正确实现方案
要实现检测字符串是否包含指定词汇,可以使用$regexMatch运算符,以下是可用写法:
写法1:多正则分开判断(方便单独调整每个关键词的匹配规则)
db.Tweets.aggregate([ { $project: { _id: 0, tweet: 1, convert: { $cond: { if: { $or: [ { $regexMatch: { input: "$tweet", regex: /Tesla/i } }, { $regexMatch: { input: "$tweet", regex: /rocket/i } }, { $regexMatch: { input: "$tweet", regex: /coffee/i } } ] }, then: "contains words", else: "does not contain words" } }, // 检测是否不存在指定词汇的布尔值字段 notContainWords: { $not: { $or: [ { $regexMatch: { input: "$tweet", regex: /Tesla/i } }, { $regexMatch: { input: "$tweet", regex: /rocket/i } }, { $regexMatch: { input: "$tweet", regex: /coffee/i } } ] } } } } ])
写法2:合并正则简化代码(关键词较多时更简洁)
db.Tweets.aggregate([ { $project: { _id: 0, tweet: 1, convert: { $cond: { if: { $regexMatch: { input: "$tweet", regex: /Tesla|rocket|coffee/i } }, then: "contains words", else: "does not contain words" } } } } ])
内容的提问来源于stack exchange,提问作者Christopher
相关产品推荐
相关产品推荐

