如何在numpy数组中重复行以高效生成模拟数据集?
用Numpy批量重复行生成模拟数据
你可以直接用Numpy的np.repeat()函数批量重复指定行,彻底替代手动复制粘贴的低效方式,完全适配大规模数据集生成需求。
针对你场景的简化代码
先定义基础的行数据,再指定每行的重复次数即可:
import numpy as np # 定义两类基础投票偏好行 base_rows = np.array([ ['Democrat', 'Republican', 'Third'], ['Democrat', 'Third', 'Republican'] ]) # 指定每行重复次数:第一行重复5次,第二行重复4次 repeat_counts = [5, 4] # 生成最终数组 voters = np.repeat(base_rows, repeat_counts, axis=0)
扩展到数千行规模
如果需要生成数千行数据,只需调整repeat_counts的数值就行,比如要生成6000行第一类、4000行第二类:
repeat_counts = [6000, 4000] voters = np.repeat(base_rows, repeat_counts, axis=0)
随机化重复(可选)
要是需要随机生成重复次数或随机抽取行,可以结合np.random模块实现:
# 随机生成每行的重复次数(范围1000到2000) random_counts = np.random.randint(1000, 2001, size=base_rows.shape[0]) voters = np.repeat(base_rows, random_counts, axis=0)
内容的提问来源于stack exchange,提问作者rulesforpower
相关产品推荐
相关产品推荐

