如何从相似字符串列表中提取差异最大的X个字符串?
从字符串数组中选出X个差异最大的字符串(基于similar_text()实现)
嘿,这个需求太贴合SEO场景了——毕竟重复内容会拉低页面权重,得尽可能选出差异最大的内容片段。我来一步步教你怎么用PHP的similar_text()实现这个功能。
核心思路
要选出X个差异最大的字符串,本质上是要找到一组字符串,它们两两之间的平均相似度最低。我们可以借助similar_text()计算每对字符串的相似度,再通过贪心算法逐步筛选出符合要求的集合。
步骤1:计算所有字符串对的相似度
similar_text()函数可以返回两个字符串的相似字符数,还能通过第三个参数获取相似度百分比(0%完全不同,100%完全相同)。我们先写一个函数,生成所有字符串两两之间的相似度矩阵:
function generateSimilarityMatrix($strings) { $count = count($strings); $matrix = array_fill(0, $count, array_fill(0, $count, 0)); for ($i = 0; $i < $count; $i++) { for ($j = $i; $j < $count; $j++) { if ($i === $j) { $matrix[$i][$j] = 100; // 自己和自己相似度100% continue; } similar_text($strings[$i], $strings[$j], $percent); $matrix[$i][$j] = $percent; $matrix[$j][$i] = $percent; // 相似度是双向的 } } return $matrix; }
步骤2:用贪心算法筛选差异最大的X个字符串
贪心算法的逻辑很简单:先选差异最大的两个字符串,之后每次加入一个和已选集合平均相似度最低的字符串,直到凑够X个。这种方法实现起来简单,而且在大部分场景下效果不错:
function selectMostDifferentStrings($strings, $X) { $count = count($strings); // 边界情况处理 if ($X <= 0) return []; if ($X >= $count) return $strings; if ($X === 1) return [$strings[array_rand($strings)]]; // 生成相似度矩阵 $similarityMatrix = generateSimilarityMatrix($strings); // 第一步:找到相似度最低的一对字符串 $minSimilarity = 100; $pair = [0, 1]; for ($i = 0; $i < $count; $i++) { for ($j = $i + 1; $j < $count; $j++) { if ($similarityMatrix[$i][$j] < $minSimilarity) { $minSimilarity = $similarityMatrix[$i][$j]; $pair = [$i, $j]; } } } // 初始化结果集合 $selectedIndices = [$pair[0], $pair[1]]; $selectedStrings = [$strings[$pair[0]], $strings[$pair[1]]]; // 逐步添加剩余的字符串 while (count($selectedIndices) < $X) { $lowestAvgSimilarity = 100; $bestIndex = -1; // 遍历未选中的每个字符串,计算它和已选集合的平均相似度 for ($i = 0; $i < $count; $i++) { if (in_array($i, $selectedIndices)) continue; $totalSimilarity = 0; foreach ($selectedIndices as $selectedIdx) { $totalSimilarity += $similarityMatrix[$i][$selectedIdx]; } $avgSimilarity = $totalSimilarity / count($selectedIndices); // 选平均相似度最低的那个 if ($avgSimilarity < $lowestAvgSimilarity) { $lowestAvgSimilarity = $avgSimilarity; $bestIndex = $i; } } // 加入结果集合 $selectedIndices[] = $bestIndex; $selectedStrings[] = $strings[$bestIndex]; } return $selectedStrings; }
测试示例
用你给出的数组测试一下:
$str = array( 'monkey eat a banana', 'dog eat a banana', 'cat devour an apple', 'cat dine a coco' ); // 提取3个差异最大的字符串 $result = selectMostDifferentStrings($str, 3); print_r($result);
运行结果应该和你预期的一致:
Array ( [0] => monkey eat a banana [1] => cat dine a coco [2] => cat devour an apple )
优化小建议
- 如果你的字符串数组特别大(比如上千条),两两计算相似度会有点慢。可以先做简单聚类,把相似的字符串归为一组,再从每组选一个代表,这样能减少计算量。
- 除了
similar_text(),也可以试试levenshtein()(编辑距离),它从字符修改次数的角度衡量差异,有时候会更适合长文本的场景。
内容的提问来源于stack exchange,提问作者JojoLapin45
相关产品推荐
相关产品推荐

