如何去除二维数组中的完全重复行并统计唯一行的出现次数?
处理二维数组重复行并统计出现次数
针对50万条数据的场景,必须用高效的O(n)复杂度方案,避免嵌套循环拖慢性能。核心思路是把每行转成唯一标识作为数组键,用键来统计重复次数,最后再还原行数据并添加统计值。
实现代码
$input = [ ['manufacturer' => 'KInd', 'brand' => 'ABC', 'used' => 'true'], ['manufacturer' => 'KInd', 'brand' => 'ABC', 'used' => 'true'], ['manufacturer' => 'KInd', 'brand' => 'ABC', 'used' => 'false'], ]; $counts = []; // 第一步:统计每行出现次数 foreach ($input as $row) { // 用json_encode生成唯一键(关联数组键顺序一致时可用,若键顺序可能乱,先ksort($row)) $key = json_encode($row); $counts[$key] = isset($counts[$key]) ? $counts[$key] + 1 : 1; } // 第二步:生成带count的结果数组 $result = []; foreach ($counts as $key => $count) { $row = json_decode($key, true); $row['count'] = $count; $result[] = $row; } print_r($result);
关键细节
- 若数组行中键的顺序可能不一致(比如同一行数据但键排列顺序不同),需要在生成键之前先对行的键排序:
ksort($row); // 先排序键,确保相同内容的行生成相同的key $key = json_encode($row); - 用
json_encode比serialize更轻量,处理大数据时性能更优;如果是纯索引数组,也可以用implode(',', $row)生成键,但关联数组必须用序列化方式。 - 整个过程仅需两次线性循环,50万条数据可快速处理,不会出现性能瓶颈。
输出结果与需求一致:
Array ( [0] => Array ( [manufacturer] => KInd [brand] => ABC [used] => true [count] => 2 ) [1] => Array ( [manufacturer] => KInd [brand] => ABC [used] => false [count] => 1 ) )
内容的提问来源于stack exchange,提问作者Fasna
相关产品推荐
相关产品推荐

