Node.js中FAQ相似问题自动聚类方案(无需指定聚类数)
无监督自动聚类FAQ相似问题的解决方案
我明白你在开发FAQ系统时遇到的痛点——手动指定聚类组数不仅麻烦,还经常得到不符合预期的结果,想要让算法自动识别合理的相似问题分组对吧?这里有几个适合的npm包,能帮你实现无监督自动聚类,我给你详细讲讲用法:
1. 使用natural实现层次聚类
natural是一个功能丰富的NLP工具包,自带层次聚类功能,不需要预先指定聚类数量,它会根据相似度自动构建聚类树,你可以通过设置阈值来得到最终的分组。
安装
npm install natural
示例代码
结合你的FAQ场景,我们可以直接基于问题文本计算相似度(也可以结合你现有的tags逻辑):
const natural = require('natural'); const { JaccardDistance, HierarchicalClustering } = natural; // 你的示例问题数据 const questions = [ "Tell me about the pricing of your product ?", "Can I talk to your agent ?", "Hi", "Hi Friend", "Hi Good Morning", "How much will it cost me ?" ]; // 预处理:转小写、去除标点 const preprocess = (text) => text.toLowerCase().replace(/[?]/g, '').trim(); const processedQuestions = questions.map(q => preprocess(q)); // 计算Jaccard相似度(也可以用余弦相似度等更精准的算法) const distanceFunc = (a, b) => { const jaccard = new JaccardDistance(); return jaccard.distance(a.split(' '), b.split(' ')); }; // 初始化层次聚类,无需指定组数 const clusterer = new HierarchicalClustering({ distance: distanceFunc, linkage: 'average' // 可选链接方式:single、complete、average等 }); // 执行聚类 const clusters = clusterer.cluster(processedQuestions); // 提取最终分组 const getGroups = (clusterTree, items) => { if (clusterTree.left && clusterTree.right) { return [...getGroups(clusterTree.left, items), ...getGroups(clusterTree.right, items)]; } else { return [[items[clusterTree]]]; } }; const finalGroups = getGroups(clusters, questions); console.log("自动聚类结果:", finalGroups);
这个代码会自动把相似的问候语、价格相关问题、代理咨询问题分成合理的三组,完全不需要你手动指定数量。
2. 使用clusterfck的DBSCAN聚类
DBSCAN是基于密度的聚类算法,完全不需要指定聚类数量,它会自动识别高密度区域作为簇,特别适合处理你的FAQ问题(比如问候语是一个高密度簇,价格问题是另一个)。
安装
npm install clusterfck
示例代码
const clusterfck = require('clusterfck'); // 问题数据 const questions = [ "Tell me about the pricing of your product ?", "Can I talk to your agent ?", "Hi", "Hi Friend", "Hi Good Morning", "How much will it cost me ?" ]; // 预处理文本 const preprocess = (text) => text.toLowerCase().replace(/[?]/g, '').trim(); const processed = questions.map(q => preprocess(q)); // 计算文本相似度(你也可以替换成基于tags的相似度逻辑) const similarity = (a, b) => { const aWords = new Set(a.split(' ')); const bWords = new Set(b.split(' ')); const intersection = [...aWords].filter(w => bWords.has(w)).length; const union = aWords.size + bWords.size - intersection; return union === 0 ? 1 : intersection / union; }; // 转换为距离值(DBSCAN用距离,1 - 相似度) const distanceMatrix = processed.map((a, i) => processed.map((b, j) => i === j ? 0 : 1 - similarity(a, b)) ); // 初始化DBSCAN,设置距离阈值和最小密度点数 const dbscan = new clusterfck.DBSCAN(); // eps=0.3表示相似度≥0.7的归为一类,minPoints=2表示至少2个点成簇 const clusters = dbscan.run(distanceMatrix, 0.3, 2); // 映射回原始问题 const finalGroups = {}; clusters.forEach((label, index) => { if (!finalGroups[label]) finalGroups[label] = []; finalGroups[label].push(questions[index]); }); console.log("DBSCAN自动聚类结果:", Object.values(finalGroups));
你可以调整eps和minPoints优化效果,比如把minPoints设为1,单独的代理咨询问题也会成为一个独立簇。
3. 使用ml-clustering的自动K-Means(肘部法则)
如果你还是倾向于用K-Means但不想手动指定k值,可以用ml-clustering的肘部法则自动选择最优的聚类数量,再进行聚类。
安装
npm install ml-clustering
示例代码
const { KMeans } = require('ml-clustering'); const natural = require('natural'); const questions = [ "Tell me about the pricing of your product ?", "Can I talk to your agent ?", "Hi", "Hi Friend", "Hi Good Morning", "How much will it cost me ?" ]; // 预处理文本 const preprocess = (text) => text.toLowerCase().replace(/[?]/g, '').trim(); // 把文本转换为TF-IDF向量 const tfidf = new natural.TfIdf(); questions.forEach(q => tfidf.addDocument(preprocess(q))); const vectors = questions.map(q => { const vec = []; tfidf.listTerms(questions.indexOf(q)).forEach(item => vec.push(item.tfidf)); return vec; }); // 肘部法则寻找最优k值 const findOptimalK = (vectors, maxK) => { const inertias = []; for (let k = 1; k <= maxK; k++) { const kmeans = new KMeans({ k }); kmeans.cluster(vectors); inertias.push(kmeans.inertia); } // 找到惯性下降最快的点(肘部点) const differences = inertias.slice(0, -1).map((val, i) => val - inertias[i+1]); const optimalK = differences.indexOf(Math.max(...differences)) + 2; return optimalK; }; const optimalK = findOptimalK(vectors, 5); console.log("自动选择的最优聚类数:", optimalK); // 用最优k值执行聚类 const kmeans = new KMeans({ k: optimalK }); const clusters = kmeans.cluster(vectors); // 映射回原始问题 const finalGroups = {}; clusters.assignments.forEach((label, index) => { if (!finalGroups[label]) finalGroups[label] = []; finalGroups[label].push(questions[index]); }); console.log("自动K-Means聚类结果:", Object.values(finalGroups));
这个方法会自动计算出适合的聚类数量,比如你的示例中会得到k=3,正好符合合理的分组逻辑。
内容的提问来源于stack exchange,提问作者Keval Bhogayata
相关产品推荐
相关产品推荐

