基于Numpy/PyTorch的1D数组特殊索引:按x的ID保留y首个1
向量化实现方案
核心逻辑
利用x为连续递增ID序列的特性,只需筛选出每个ID分组内y值第一个为1的位置,其余位置置0即可,全程无Python层面循环,性能远高于原生循环实现。
NumPy 实现
import numpy as np x = np.array([0, 0, 0, 1, 1, 2, 2, 2, 2]) y = np.array([1, 1, 0, 1, 1, 0, 0, 1, 0]) z = np.zeros_like(y) # 提取所有y值为1的索引 y1_idx = np.where(y == 1)[0] # 提取这些索引对应的x值 x_at_y1 = x[y1_idx] # 标记每个ID在y=1的位置中首次出现的位置 first_occur_mask = np.r_[True, x_at_y1[1:] != x_at_y1[:-1]] # 仅保留首次出现的1 z[y1_idx[first_occur_mask]] = 1
PyTorch 实现
import torch x = torch.tensor([0, 0, 0, 1, 1, 2, 2, 2, 2]) y = torch.tensor([1, 1, 0, 1, 1, 0, 0, 1, 0]) z = torch.zeros_like(y) y1_idx = torch.where(y == 1)[0] x_at_y1 = x[y1_idx] first_occur_mask = torch.cat([torch.tensor([True], device=x.device), x_at_y1[1:] != x_at_y1[:-1]]) z[y1_idx[first_occur_mask]] = 1
效果验证
以示例输入运行上述代码,输出z为[1 0 0 1 0 0 0 1 0],完全符合需求。
内容的提问来源于stack exchange,提问作者tphillips
相关产品推荐
相关产品推荐

