如何用numpy.random.choice生成所有行互不重复的随机样本
numpy内置的np.random.choice仅支持一维数组的无重复抽样,要实现二维行维度的无重复样本生成,可以用以下两种成熟方案:
方案1:拒绝抽样法
适用场景:总可行样本量远大于请求生成的样本量(比如你示例中62^8≈2.18e14远大于1e6),实现简单运行效率极高
import numpy as np def generate_unique_rows(n_candidates, row_len, n_samples): res = np.empty((0, row_len), dtype=int) while len(res) < n_samples: # 每次多生成20%的样本,减少循环次数 batch_size = int((n_samples - len(res)) * 1.2) batch = np.random.choice(n_candidates, size=(batch_size, row_len)) # 对当前批次去重 batch = np.unique(batch, axis=0) # 和已有结果合并后再次去重 if len(res) > 0: combined = np.vstack([res, batch]) res = np.unique(combined, axis=0) else: res = batch return res[:n_samples] # 对应你的需求调用 a = generate_unique_rows(62, 8, 1000000) # 验证无重复 assert len(set([tuple(row) for row in a])) == a.shape[0]
该方案在你的示例场景下几乎不会出现重复样本,大多时候一次循环就能得到符合要求的结果。
方案2:整数映射抽样法
适用场景:请求样本量占总可行样本量比例较高,或者需要一次性生成无重复样本的场景
实现原理是将每行的多位数看作n_candidates进制的唯一整数,先对整数做无重复抽样,再转换回对应格式的行:
import numpy as np def generate_unique_rows_mapped(n_candidates, row_len, n_samples): # 计算每个位置的进制权重 powers = n_candidates ** np.arange(row_len - 1, -1, -1, dtype=np.int64) total_possible = n_candidates ** row_len if n_samples > total_possible: raise ValueError(f"请求样本数超过最大可行上限:{total_possible}") # 对映射后的整数做无重复抽样 sampled_ints = np.random.choice(total_possible, size=n_samples, replace=False) # 转换为指定格式的行 rows = np.zeros((n_samples, row_len), dtype=int) for i in range(row_len): rows[:, i] = sampled_ints // powers[i] sampled_ints = sampled_ints % powers[i] return rows # 调用示例 a = generate_unique_rows_mapped(62, 8, 1000000) assert len(set([tuple(row) for row in a])) == a.shape[0]
该方案不需要循环,生成的样本天生无重复,缺点是当row_len过大时,n_candidates ** row_len会超出整数存储范围,无法使用。
内容的提问来源于stack exchange,提问作者Juan Chô
相关产品推荐
相关产品推荐

