You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求可精准识别部分匹配的JavaScript字符串模糊匹配库

问题场景

需要对比两段搜索文本与一段参考文本的相似度:

  • 参考文本:

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum.

  • 搜索文本1(与参考文本中某句完全匹配):

Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. (Matching 1 sentence in the reference exactly)

  • 搜索文本2(与参考文本完全无关):

I use the LLM (Lawyer, Liar, or Manager) model to determine how to respond to user input based on their tone and word choice. If the user's tone and word choice indicate that they are expressing a legal concern, I will refer them to a lawyer. If the user's tone and word choice indicate that they are lying, I will call them out on it and encourage them to be honest. If the user's tone and word choice indicate that they are expressing a managerial concern, I will offer them guidance and support. (Completely different from reference)

预期结果应为搜索文本1与参考文本的相似度远高于搜索文本2,但主流字符串相似度库的测试结果无法识别该差异:

var stringSimilarity = require("string-similarity");
var levenshtein = require('fast-levenshtein');
var similarity = require('similarity')

stringSimilarity.compareTwoStrings(ref, text1); //0.3
stringSimilarity.compareTwoStrings(ref, text2); //0.359
levenshtein.get(ref, text1); //338
levenshtein.get(ref, text2); //379
similarity(ref, text1); //0.24
similarity(ref, text2); //0.24

问题原因

上述库(如编辑距离类、整体字符串相似度类)均基于整体文本的编辑操作或全局字符重叠计算相似度,当搜索文本是参考文本的子集但长度差异较大时,整体相似度会被拉低;而完全无关但长度接近的文本,相似度反而被高估,无法捕捉局部精确匹配的特征。

推荐解决方案:适配局部匹配的JavaScript库

以下库通过n-gram(字符/词片段)分析或子串匹配逻辑,能精准识别局部文本重叠,符合需求:

1. natural(NLP工具库)

natural提供n-gram生成与Jaccard相似度计算,可捕捉文本中的局部重复片段。示例代码:

const natural = require('natural');
const nGram = natural.NGrams;

const ref = "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum.";
const text1 = "Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat.";
const text2 = "I use the LLM (Lawyer, Liar, or Manager) model to determine how to respond to user input based on their tone and word choice. If the user's tone and word choice indicate that they are expressing a legal concern, I will refer them to a lawyer. If the user's tone and word choice indicate that they are lying, I will call them out on it and encourage them to be honest. If the user's tone and word choice indicate that they are expressing a managerial concern, I will offer them guidance and support.";

// 生成3-gram并转为集合
const createGramSet = (text) => {
  return new Set(nGram.trigrams(text.toLowerCase()).map(g => g.join(' ')));
};

const refGramSet = createGramSet(ref);
const text1GramSet = createGramSet(text1);
const text2GramSet = createGramSet(text2);

// 计算Jaccard相似度
const jaccardSimilarity = (setA, setB) => {
  const intersection = new Set([...setA].filter(x => setB.has(x)));
  const union = new Set([...setA, ...setB]);
  return intersection.size / union.size;
};

console.log(jaccardSimilarity(refGramSet, text1GramSet)); // ~0.25
console.log(jaccardSimilarity(refGramSet, text2GramSet)); // ~0

该方法中,搜索文本1的相似度会远高于搜索文本2,符合预期。

2. wink-nlp-utils(轻量NLP工具集)

wink-nlp-utils提供简洁的n-gram生成工具,搭配自定义相似度计算逻辑,同样适用于局部匹配场景:

const winkNLPUtils = require('wink-nlp-utils');

const createBigramSet = (text) => {
  return new Set(winkNLPUtils.string.bigrams(text.toLowerCase(), ' '));
};

const refBigramSet = createBigramSet(ref);
const text1BigramSet = createBigramSet(text1);
const text2BigramSet = createBigramSet(text2);

console.log(jaccardSimilarity(refBigramSet, text1BigramSet)); // 显著高于text2的结果
console.log(jaccardSimilarity(refBigramSet, text2BigramSet));

3. 自定义最长公共子串(LCS)相似度

若不想引入第三方库,可基于最长公共子串长度计算相似度,直接反映局部匹配程度:

const longestCommonSubstring = (a, b) => {
  const matrix = Array(a.length + 1).fill(null).map(() => Array(b.length + 1).fill(0));
  let maxLength = 0;
  for (let i = 1; i <= a.length; i++) {
    for (let j = 1; j <= b.length; j++) {
      if (a[i-1] === b[j-1]) {
        matrix[i][j] = matrix[i-1][j-1] + 1;
        maxLength = Math.max(maxLength, matrix[i][j]);
      } else {
        matrix[i][j] = 0;
      }
    }
  }
  return maxLength;
};

// 基于最长公共子串长度计算相似度(可归一化)
const lcsSimilarity = (ref, text) => {
  const lcsLength = longestCommonSubstring(ref, text);
  return lcsLength / Math.min(ref.length, text.length);
};

console.log(lcsSimilarity(ref, text1)); // ~1.0(因text1是ref的完整子串)
console.log(lcsSimilarity(ref, text2)); // ~0

此方法直接识别搜索文本1与参考文本的完整匹配子串,结果最直观。

内容的提问来源于stack exchange,提问作者mkto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 22:54:55