You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDL多维数组索引获取:如何高效获取元素出现的索引?

解决PDL中获取数组元素对应索引的高效方案

核心思路:排序分组法

通过将原数组与索引绑定后排序,利用RLE(游程编码)识别相同元素的连续区间,一次性提取所有唯一元素的索引,避免循环调用which带来的性能损耗,适合大数据量场景。

具体实现代码

use PDL;

# 初始化数据
my $a = pdl(3,4,6,3,2,3,5,6,1,7);
my $indices = sequence($a->dim(0)); # 生成与$a长度一致的索引序列,等价于示例中的$aa

# 将原数组与索引堆叠,按原数组的值排序
my $sorted_pair = cat($a, $indices)->qsortvec;

# 拆分排序后的数值和索引
my $sorted_vals = $sorted_pair->slice(':,(0)');
my $sorted_indices = $sorted_pair->slice(':,(1)');

# 用RLE获取唯一值及对应出现次数
my ($unique_vals, $counts) = rle($sorted_vals);

# 计算每个唯一值分组的起始/结束位置
my $starts = zeroes($counts->dim(0));
$starts->slice(1:) .= $counts->cumusum->slice(0:-2);
my $ends = $starts + $counts - 1;

# 构建唯一值到对应索引的映射
my %value_to_indices;
for my $i (0 .. $unique_vals->dim(0)-1) {
    my $start = $starts->at($i);
    my $end = $ends->at($i);
    $value_to_indices{$unique_vals->at($i)} = $sorted_indices->slice("$start:$end")->unpdl;
}

# 验证结果:打印元素3对应的索引
print "元素3的索引:", join(', ', @{$value_to_indices{3}}), "\n";
# 输出:元素3的索引:0, 3, 5

备选方案:基于布尔矩阵的批量提取

如果不需要哈希映射,也可以直接利用示例中的布尔矩阵$c,批量提取每列的索引:

my $b = uniq($a);
my $c = $a->(*1) == $b;

# 提取每个唯一值对应的索引
my @all_indices;
for my $col_idx (0 .. $b->dim(0)-1) {
    push @all_indices, which($c->slice(",$col_idx"));
}

# 输出第一个唯一值(1)的索引
print "元素1的索引:", $all_indices[0], "\n";
# 输出:元素1的索引:8

注意:此方案仍需循环处理每一列,在唯一值数量极多时,性能不如排序分组法。

内容的提问来源于stack exchange,提问作者doosoonk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 12:21:23