如何在MATLAB中高效获取坐标列表中各坐标的首尾出现索引
Hey there! Since you're dealing with 30 million coordinates, ditching that loop is absolutely critical for performance—great job getting the last occurrence vectorized already. Let's tackle the first occurrence problem with a few efficient vectorized methods that play nicely with large datasets.
Method 1: Accumarray (Matching Your Existing Matrix-Based Workflow)
This aligns perfectly with how you handled Mlast, but instead of overwriting values to keep the latest index, we'll use accumarray to grab the minimum index (which is the first occurrence) for each unique coordinate pair.
Here's the code:
x1 = min(x); y1 = min(y); x2 = max(x); y2 = max(y); ind = sub2ind([y2-y1+1, x2-x1+1], y-y1+1, x-x1+1); % Get first occurrence indices using accumarray with @min Mfirst = accumarray(ind, (1:length(x))', [], @min); % Map back to the original order to get your 'first' array first = Mfirst(ind);
Why this works:
accumarraygroups all indices that share the sameind(i.e., same coordinate pair) and applies the@minfunction to them. The smallest index in each group is exactly the first occurrence.- This uses the same linear indexing setup you already have for
Mlast, so it's a natural extension of your existing code.
Note: If your x/y values have an extremely large range (e.g., spanning millions of integers), the Mfirst matrix might eat up too much memory. In that case, use the method below.
Method 2: Unique Coordinate Pairs (Memory-Friendly for Large Value Ranges)
If your x/y coordinates don't fit neatly into a small matrix, we can use unique to directly track first occurrences of each (x,y) pair:
% Combine x and y into a matrix of coordinate pairs coords = [x(:), y(:)]; % Get unique pairs and their first occurrence indices [~, first_idx, group_id] = unique(coords, 'rows', 'first'); % Map group IDs back to first occurrence indices first = first_idx(group_id);
Why this works:
unique(..., 'rows', 'first')returnsfirst_idx(the index of the first occurrence of each unique pair) andgroup_id(a vector where each element tells which unique group the original coordinate belongs to).- Indexing
first_idxwithgroup_idgives us the first occurrence index for every original coordinate.
This method uses way less memory when your x/y values have a sparse or large range, since it only stores data for unique pairs instead of a full matrix.
Performance Note
Both methods will blow the loop out of the water for 30 million entries—MATLAB's vectorized operations are optimized for large datasets, so you'll see a massive speedup. Test both to see which fits your memory constraints best!
内容的提问来源于stack exchange,提问作者MiB_Coder

