JavaScript实现大字符串与字符串数组的3词及以上连续匹配问题
问题
我有一个包含数千个单词的大字符串,需要与数组中同样为大字符串的所有元素进行对比,找出所有3个及以上连续匹配的单词。用正则表达式实现后得到的却是空匹配数组。
示例(简化文本)
let textToCompare = "Hello there how are you doing with your life"; let textsToCompareWith= [ { id:1, text:"Hope you are doing good with your life" }, { id:2, text:"what are you doing with your life. hello there how are you" }, { id:3, text:"hello there mate" } ];
预期输出
[ {id:1, matchedText:["with your life"]}, {id:2, matchedText:["are you doing with your life","hello there how are you"]}, {id:3, matchedText:[]} ];
当前输出
[ {id:1, matchedText:[]}, {id:2, matchedText:[]}, {id:3, matchedText:[]} ];
我的代码
let regex = new RegExp("\\b" + textToCompare.split(" ").join("\\b.*\\b") + "\\b", "gi"); let output = textsToCompareWith.map(textObj => { // Match against each element in the array let matchedText = textObj?.text.match(regex); console.log(matchedText); return { id: textObj.id, matchedText: matchedText ? matchedText : [] // Return an empty array if no match is found }; }); console.log(output);
解决方案
你的正则表达式逻辑完全错误:它会要求匹配整个textToCompare的单词顺序,中间允许插入任意字符,而不是从目标文本中找出原字符串里连续3个及以上单词的片段。要实现需求,需要换思路:
核心思路
- 先从原字符串中生成所有长度≥3的连续单词组合(比如
"Hello there how"、"are you doing with your life"等) - 对每个组合生成正则,在目标文本中匹配完整单词序列,最后去重整理结果
修正后的代码
const textToCompare = "Hello there how are you doing with your life"; const sourceWords = textToCompare.split(" "); // 生成所有长度≥3的连续单词组合 const candidatePhrases = []; for (let i = 0; i <= sourceWords.length - 3; i++) { for (let j = i + 2; j < sourceWords.length; j++) { candidatePhrases.push(sourceWords.slice(i, j + 1).join(" ")); } } // 去重,减少重复匹配次数 const uniquePhrases = [...new Set(candidatePhrases)]; const textsToCompareWith= [ { id:1, text:"Hope you are doing good with your life" }, { id:2, text:"what are you doing with your life. hello there how are you" }, { id:3, text:"hello there mate" } ]; const output = textsToCompareWith.map(textObj => { const matched = []; uniquePhrases.forEach(phrase => { // 用单词边界确保完整单词匹配,不区分大小写 const regex = new RegExp(`\\b${phrase}\\b`, "gi"); const matches = textObj.text.match(regex); if (matches) { // 避免同一文本重复添加相同匹配项 matches.forEach(match => { const lowerMatch = match.toLowerCase(); if (!matched.some(item => item.toLowerCase() === lowerMatch)) { matched.push(match); } }); } }); // 按长度降序排列,优先保留更长的匹配片段(可选) matched.sort((a, b) => b.split(" ").length - a.split(" ").length); return { id: textObj.id, matchedText: matched }; }); console.log(output);
代码说明
- 生成候选片段:遍历原字符串的单词数组,生成所有长度从3到总单词数的连续组合,覆盖所有可能的匹配项
- 去重处理:避免重复生成相同短语,减少不必要的正则匹配
- 正则匹配:用
\b包裹短语,确保匹配完整单词,不会出现you匹配your的情况 - 结果去重:同一文本中可能多次出现相同匹配片段,去重后返回更简洁的结果
性能优化提示
如果原字符串包含上万级别的单词,预生成所有候选片段会占用大量内存。此时可以改用滑动窗口的方式,直接在目标文本中拆分单词,然后与原字符串的单词数组做连续匹配对比,避免预生成所有组合。
内容的提问来源于stack exchange,提问作者BSA
相关产品推荐
相关产品推荐

