You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从相似字符串列表中提取差异最大的X个字符串?

从字符串数组中选出X个差异最大的字符串(基于similar_text()实现)

嘿,这个需求太贴合SEO场景了——毕竟重复内容会拉低页面权重,得尽可能选出差异最大的内容片段。我来一步步教你怎么用PHP的similar_text()实现这个功能。

核心思路

要选出X个差异最大的字符串,本质上是要找到一组字符串,它们两两之间的平均相似度最低。我们可以借助similar_text()计算每对字符串的相似度,再通过贪心算法逐步筛选出符合要求的集合。

步骤1:计算所有字符串对的相似度

similar_text()函数可以返回两个字符串的相似字符数,还能通过第三个参数获取相似度百分比(0%完全不同,100%完全相同)。我们先写一个函数,生成所有字符串两两之间的相似度矩阵:

function generateSimilarityMatrix($strings) {
    $count = count($strings);
    $matrix = array_fill(0, $count, array_fill(0, $count, 0));
    
    for ($i = 0; $i < $count; $i++) {
        for ($j = $i; $j < $count; $j++) {
            if ($i === $j) {
                $matrix[$i][$j] = 100; // 自己和自己相似度100%
                continue;
            }
            similar_text($strings[$i], $strings[$j], $percent);
            $matrix[$i][$j] = $percent;
            $matrix[$j][$i] = $percent; // 相似度是双向的
        }
    }
    
    return $matrix;
}

步骤2:用贪心算法筛选差异最大的X个字符串

贪心算法的逻辑很简单:先选差异最大的两个字符串,之后每次加入一个和已选集合平均相似度最低的字符串,直到凑够X个。这种方法实现起来简单,而且在大部分场景下效果不错:

function selectMostDifferentStrings($strings, $X) {
    $count = count($strings);
    
    // 边界情况处理
    if ($X <= 0) return [];
    if ($X >= $count) return $strings;
    if ($X === 1) return [$strings[array_rand($strings)]];
    
    // 生成相似度矩阵
    $similarityMatrix = generateSimilarityMatrix($strings);
    
    // 第一步:找到相似度最低的一对字符串
    $minSimilarity = 100;
    $pair = [0, 1];
    for ($i = 0; $i < $count; $i++) {
        for ($j = $i + 1; $j < $count; $j++) {
            if ($similarityMatrix[$i][$j] < $minSimilarity) {
                $minSimilarity = $similarityMatrix[$i][$j];
                $pair = [$i, $j];
            }
        }
    }
    
    // 初始化结果集合
    $selectedIndices = [$pair[0], $pair[1]];
    $selectedStrings = [$strings[$pair[0]], $strings[$pair[1]]];
    
    // 逐步添加剩余的字符串
    while (count($selectedIndices) < $X) {
        $lowestAvgSimilarity = 100;
        $bestIndex = -1;
        
        // 遍历未选中的每个字符串,计算它和已选集合的平均相似度
        for ($i = 0; $i < $count; $i++) {
            if (in_array($i, $selectedIndices)) continue;
            
            $totalSimilarity = 0;
            foreach ($selectedIndices as $selectedIdx) {
                $totalSimilarity += $similarityMatrix[$i][$selectedIdx];
            }
            $avgSimilarity = $totalSimilarity / count($selectedIndices);
            
            // 选平均相似度最低的那个
            if ($avgSimilarity < $lowestAvgSimilarity) {
                $lowestAvgSimilarity = $avgSimilarity;
                $bestIndex = $i;
            }
        }
        
        // 加入结果集合
        $selectedIndices[] = $bestIndex;
        $selectedStrings[] = $strings[$bestIndex];
    }
    
    return $selectedStrings;
}

测试示例

用你给出的数组测试一下:

$str = array(
    'monkey eat a banana',
    'dog eat a banana',
    'cat devour an apple',
    'cat dine a coco'
);

// 提取3个差异最大的字符串
$result = selectMostDifferentStrings($str, 3);
print_r($result);

运行结果应该和你预期的一致:

Array
(
    [0] => monkey eat a banana
    [1] => cat dine a coco
    [2] => cat devour an apple
)

优化小建议

  • 如果你的字符串数组特别大(比如上千条),两两计算相似度会有点慢。可以先做简单聚类,把相似的字符串归为一组,再从每组选一个代表,这样能减少计算量。
  • 除了similar_text(),也可以试试levenshtein()(编辑距离),它从字符修改次数的角度衡量差异,有时候会更适合长文本的场景。

内容的提问来源于stack exchange,提问作者JojoLapin45

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:35:16