You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JavaScript实现大字符串与字符串数组的3词及以上连续匹配问题

问题

我有一个包含数千个单词的大字符串,需要与数组中同样为大字符串的所有元素进行对比,找出所有3个及以上连续匹配的单词。用正则表达式实现后得到的却是空匹配数组。

示例(简化文本)

let textToCompare = "Hello there how are you doing with your life";

let textsToCompareWith= [
  { id:1, text:"Hope you are doing good with your life" },
  { id:2, text:"what are you doing with your life. hello there how are you" },
  { id:3, text:"hello there mate" }
];

预期输出

[
  {id:1, matchedText:["with your life"]}, 
  {id:2, matchedText:["are you doing with your life","hello there how are you"]},
  {id:3, matchedText:[]}
];

当前输出

[
  {id:1, matchedText:[]}, 
  {id:2, matchedText:[]},
  {id:3, matchedText:[]}
];

我的代码

let regex = new RegExp("\\b" + textToCompare.split(" ").join("\\b.*\\b") + "\\b", "gi");

let output = textsToCompareWith.map(textObj => {
  // Match against each element in the array
  let matchedText = textObj?.text.match(regex);
  console.log(matchedText);
  return {
    id: textObj.id,
    matchedText: matchedText ? matchedText : [] // Return an empty array if no match is found
  };
});
                                  
console.log(output);

解决方案

你的正则表达式逻辑完全错误:它会要求匹配整个textToCompare的单词顺序,中间允许插入任意字符,而不是从目标文本中找出原字符串里连续3个及以上单词的片段。要实现需求,需要换思路:

核心思路

  1. 先从原字符串中生成所有长度≥3的连续单词组合(比如"Hello there how"、"are you doing with your life"等)
  2. 对每个组合生成正则,在目标文本中匹配完整单词序列,最后去重整理结果

修正后的代码

const textToCompare = "Hello there how are you doing with your life";
const sourceWords = textToCompare.split(" ");

// 生成所有长度≥3的连续单词组合
const candidatePhrases = [];
for (let i = 0; i <= sourceWords.length - 3; i++) {
  for (let j = i + 2; j < sourceWords.length; j++) {
    candidatePhrases.push(sourceWords.slice(i, j + 1).join(" "));
  }
}

// 去重,减少重复匹配次数
const uniquePhrases = [...new Set(candidatePhrases)];

const textsToCompareWith= [
  { id:1, text:"Hope you are doing good with your life" },
  { id:2, text:"what are you doing with your life. hello there how are you" },
  { id:3, text:"hello there mate" }
];

const output = textsToCompareWith.map(textObj => {
  const matched = [];
  uniquePhrases.forEach(phrase => {
    // 用单词边界确保完整单词匹配,不区分大小写
    const regex = new RegExp(`\\b${phrase}\\b`, "gi");
    const matches = textObj.text.match(regex);
    if (matches) {
      // 避免同一文本重复添加相同匹配项
      matches.forEach(match => {
        const lowerMatch = match.toLowerCase();
        if (!matched.some(item => item.toLowerCase() === lowerMatch)) {
          matched.push(match);
        }
      });
    }
  });
  // 按长度降序排列,优先保留更长的匹配片段(可选)
  matched.sort((a, b) => b.split(" ").length - a.split(" ").length);
  return {
    id: textObj.id,
    matchedText: matched
  };
});

console.log(output);

代码说明

  • 生成候选片段:遍历原字符串的单词数组,生成所有长度从3到总单词数的连续组合,覆盖所有可能的匹配项
  • 去重处理:避免重复生成相同短语,减少不必要的正则匹配
  • 正则匹配:用\b包裹短语,确保匹配完整单词,不会出现you匹配your的情况
  • 结果去重:同一文本中可能多次出现相同匹配片段,去重后返回更简洁的结果

性能优化提示

如果原字符串包含上万级别的单词,预生成所有候选片段会占用大量内存。此时可以改用滑动窗口的方式,直接在目标文本中拆分单词,然后与原字符串的单词数组做连续匹配对比,避免预生成所有组合。

内容的提问来源于stack exchange,提问作者BSA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 03:07:03