You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Node.js中FAQ相似问题自动聚类方案(无需指定聚类数)

无监督自动聚类FAQ相似问题的解决方案

我明白你在开发FAQ系统时遇到的痛点——手动指定聚类组数不仅麻烦,还经常得到不符合预期的结果,想要让算法自动识别合理的相似问题分组对吧?这里有几个适合的npm包,能帮你实现无监督自动聚类,我给你详细讲讲用法:

1. 使用natural实现层次聚类

natural是一个功能丰富的NLP工具包,自带层次聚类功能,不需要预先指定聚类数量,它会根据相似度自动构建聚类树,你可以通过设置阈值来得到最终的分组。

安装

npm install natural

示例代码

结合你的FAQ场景,我们可以直接基于问题文本计算相似度(也可以结合你现有的tags逻辑):

const natural = require('natural');
const { JaccardDistance, HierarchicalClustering } = natural;

// 你的示例问题数据
const questions = [
  "Tell me about the pricing of your product ?",
  "Can I talk to your agent ?",
  "Hi",
  "Hi Friend",
  "Hi Good Morning",
  "How much will it cost me ?"
];

// 预处理:转小写、去除标点
const preprocess = (text) => text.toLowerCase().replace(/[?]/g, '').trim();
const processedQuestions = questions.map(q => preprocess(q));

// 计算Jaccard相似度(也可以用余弦相似度等更精准的算法)
const distanceFunc = (a, b) => {
  const jaccard = new JaccardDistance();
  return jaccard.distance(a.split(' '), b.split(' '));
};

// 初始化层次聚类,无需指定组数
const clusterer = new HierarchicalClustering({
  distance: distanceFunc,
  linkage: 'average' // 可选链接方式:single、complete、average等
});

// 执行聚类
const clusters = clusterer.cluster(processedQuestions);

// 提取最终分组
const getGroups = (clusterTree, items) => {
  if (clusterTree.left && clusterTree.right) {
    return [...getGroups(clusterTree.left, items), ...getGroups(clusterTree.right, items)];
  } else {
    return [[items[clusterTree]]];
  }
};

const finalGroups = getGroups(clusters, questions);
console.log("自动聚类结果:", finalGroups);

这个代码会自动把相似的问候语、价格相关问题、代理咨询问题分成合理的三组,完全不需要你手动指定数量。

2. 使用clusterfck的DBSCAN聚类

DBSCAN是基于密度的聚类算法,完全不需要指定聚类数量,它会自动识别高密度区域作为簇,特别适合处理你的FAQ问题(比如问候语是一个高密度簇,价格问题是另一个)。

安装

npm install clusterfck

示例代码

const clusterfck = require('clusterfck');

// 问题数据
const questions = [
  "Tell me about the pricing of your product ?",
  "Can I talk to your agent ?",
  "Hi",
  "Hi Friend",
  "Hi Good Morning",
  "How much will it cost me ?"
];

// 预处理文本
const preprocess = (text) => text.toLowerCase().replace(/[?]/g, '').trim();
const processed = questions.map(q => preprocess(q));

// 计算文本相似度(你也可以替换成基于tags的相似度逻辑)
const similarity = (a, b) => {
  const aWords = new Set(a.split(' '));
  const bWords = new Set(b.split(' '));
  const intersection = [...aWords].filter(w => bWords.has(w)).length;
  const union = aWords.size + bWords.size - intersection;
  return union === 0 ? 1 : intersection / union;
};

// 转换为距离值(DBSCAN用距离,1 - 相似度)
const distanceMatrix = processed.map((a, i) => 
  processed.map((b, j) => i === j ? 0 : 1 - similarity(a, b))
);

// 初始化DBSCAN,设置距离阈值和最小密度点数
const dbscan = new clusterfck.DBSCAN();
// eps=0.3表示相似度≥0.7的归为一类,minPoints=2表示至少2个点成簇
const clusters = dbscan.run(distanceMatrix, 0.3, 2); 

// 映射回原始问题
const finalGroups = {};
clusters.forEach((label, index) => {
  if (!finalGroups[label]) finalGroups[label] = [];
  finalGroups[label].push(questions[index]);
});

console.log("DBSCAN自动聚类结果:", Object.values(finalGroups));

你可以调整eps和minPoints优化效果,比如把minPoints设为1,单独的代理咨询问题也会成为一个独立簇。

3. 使用ml-clustering的自动K-Means(肘部法则)

如果你还是倾向于用K-Means但不想手动指定k值,可以用ml-clustering的肘部法则自动选择最优的聚类数量,再进行聚类。

安装

npm install ml-clustering

示例代码

const { KMeans } = require('ml-clustering');
const natural = require('natural');

const questions = [
  "Tell me about the pricing of your product ?",
  "Can I talk to your agent ?",
  "Hi",
  "Hi Friend",
  "Hi Good Morning",
  "How much will it cost me ?"
];

// 预处理文本
const preprocess = (text) => text.toLowerCase().replace(/[?]/g, '').trim();

// 把文本转换为TF-IDF向量
const tfidf = new natural.TfIdf();
questions.forEach(q => tfidf.addDocument(preprocess(q)));

const vectors = questions.map(q => {
  const vec = [];
  tfidf.listTerms(questions.indexOf(q)).forEach(item => vec.push(item.tfidf));
  return vec;
});

// 肘部法则寻找最优k值
const findOptimalK = (vectors, maxK) => {
  const inertias = [];
  for (let k = 1; k <= maxK; k++) {
    const kmeans = new KMeans({ k });
    kmeans.cluster(vectors);
    inertias.push(kmeans.inertia);
  }
  // 找到惯性下降最快的点(肘部点)
  const differences = inertias.slice(0, -1).map((val, i) => val - inertias[i+1]);
  const optimalK = differences.indexOf(Math.max(...differences)) + 2;
  return optimalK;
};

const optimalK = findOptimalK(vectors, 5);
console.log("自动选择的最优聚类数:", optimalK);

// 用最优k值执行聚类
const kmeans = new KMeans({ k: optimalK });
const clusters = kmeans.cluster(vectors);

// 映射回原始问题
const finalGroups = {};
clusters.assignments.forEach((label, index) => {
  if (!finalGroups[label]) finalGroups[label] = [];
  finalGroups[label].push(questions[index]);
});

console.log("自动K-Means聚类结果:", Object.values(finalGroups));

这个方法会自动计算出适合的聚类数量,比如你的示例中会得到k=3,正好符合合理的分组逻辑。

内容的提问来源于stack exchange,提问作者Keval Bhogayata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:05:09