如何在Python中生成类sklearn.make_classification的结构化计数型组成数据?
生成带结构的计数型分类合成数据
针对你的需求,这里提供两种可行方案:一是改造sklearn.make_classification的输出以适配计数型数据,二是直接生成具备指定结构的计数数据,两种方案都能支持你提到的核心参数(类别数、特征数、信息特征数、冗余特征数)以及额外的样本总计数、稀疏度参数。
方案一:改造make_classification输出复用特征结构
利用make_classification已经成熟的特征结构生成逻辑(自动区分信息/冗余/噪声特征),将其输出的连续数据转换为带0的正整数计数数据,同时控制总计数和稀疏度:
import numpy as np from sklearn.datasets import make_classification def make_count_classification(n_samples=1000, n_features=20, n_informative=5, n_redundant=5, n_classes=2, total_count=100, sparsity=0.3): # 生成带结构的连续型分类数据 X, y = make_classification(n_samples=n_samples, n_features=n_features, n_informative=n_informative, n_redundant=n_redundant, n_classes=n_classes, random_state=42) # 转换为非负值(计数不能为负) X_nonneg = np.abs(X) # 归一化后缩放到指定总计数并取整 row_sums = X_nonneg.sum(axis=1, keepdims=True) X_normalized = X_nonneg / row_sums X_count = np.round(X_normalized * total_count).astype(int) # 修正取整导致的行和偏差 diff = total_count - X_count.sum(axis=1) for i in range(n_samples): if diff[i] != 0: # 随机选择非0特征调整计数,确保行和等于total_count valid_idx = np.where(X_count[i] > 0)[0] if len(valid_idx) == 0: valid_idx = np.random.choice(n_features, 3) adjust_idx = np.random.choice(valid_idx, abs(diff[i]), replace=True) X_count[i, adjust_idx] += 1 if diff[i] > 0 else -1 # 调整稀疏度:将低计数特征设为0 flat_values = X_count.flatten() threshold = np.percentile(flat_values, sparsity * 100) X_count[X_count <= threshold] = 0 # 处理调整后行和为0的样本 zero_row_mask = X_count.sum(axis=1) == 0 for i in np.where(zero_row_mask)[0]: X_count[i, np.random.choice(n_features, 3)] = np.random.randint(1, total_count//3, size=3) # 再次校准行和到total_count row_sums = X_count.sum(axis=1, keepdims=True) X_count = np.round(X_count / row_sums * total_count).astype(int) return X_count, y
方案说明
- 保留了
make_classification的特征结构:信息特征与标签强关联,冗余特征由信息特征线性组合生成,噪声特征无关联 - 通过绝对值转换、归一化缩放实现从连续数据到计数数据的映射
- 利用分位数阈值控制0值占比,精准匹配指定的稀疏度
方案二:直接生成结构化计数数据
如果希望更贴合计数数据的分布特性(比如泊松/负二项分布),可以直接构建具备结构的计数特征:
import numpy as np def make_structured_count_data(n_samples=1000, n_features=20, n_informative=5, n_redundant=5, n_classes=2, total_count=100, sparsity=0.3): # 生成标签 y = np.random.randint(0, n_classes, size=n_samples) # 生成信息特征:与标签关联的泊松计数 X_informative = [] for cls in range(n_classes): # 不同类别对应不同的泊松均值,强化特征与标签的关联 cls_means = np.random.randint(5, 20, size=n_informative) cls_samples = np.random.poisson(cls_means, size=(np.sum(y == cls), n_informative)) X_informative.append(cls_samples) X_informative = np.vstack(X_informative) # 生成冗余特征:基于信息特征的线性组合生成计数 X_redundant = [] for _ in range(n_redundant): weights = np.random.uniform(0.2, 1.0, size=n_informative) linear_comb = np.dot(X_informative, weights) # 转换为泊松计数 feat = np.random.poisson(linear_comb).reshape(-1, 1) X_redundant.append(feat) X_redundant = np.hstack(X_redundant) # 生成噪声特征:无关联的随机泊松计数 n_noise = n_features - n_informative - n_redundant X_noise = np.random.poisson(3, size=(n_samples, n_noise)) # 合并所有特征 X = np.hstack([X_informative, X_redundant, X_noise]) # 调整总计数 row_sums = X.sum(axis=1, keepdims=True) X_count = np.round(X / row_sums * total_count).astype(int) # 调整稀疏度 flat_values = X_count.flatten() threshold = np.percentile(flat_values, sparsity * 100) X_count[X_count <= threshold] = 0 # 处理行和为0的样本 zero_row_mask = X_count.sum(axis=1) == 0 for i in np.where(zero_row_mask)[0]: X_count[i, np.random.choice(n_features, 3)] = np.random.randint(1, total_count//3, size=3) row_sums = X_count.sum(axis=1, keepdims=True) X_count = np.round(X_count / row_sums * total_count).astype(int) return X_count, y
方案说明
- 信息特征直接基于标签生成,不同类别使用不同的泊松均值,确保特征与标签的关联性
- 冗余特征由信息特征线性组合后转换为计数,符合实际场景中冗余特征的生成逻辑
- 噪声特征采用独立泊松分布生成,模拟无意义的计数特征
内容的提问来源于stack exchange,提问作者O.rka
相关产品推荐
相关产品推荐

