You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中生成类sklearn.make_classification的结构化计数型组成数据?

生成带结构的计数型分类合成数据

针对你的需求,这里提供两种可行方案:一是改造sklearn.make_classification的输出以适配计数型数据,二是直接生成具备指定结构的计数数据,两种方案都能支持你提到的核心参数(类别数、特征数、信息特征数、冗余特征数)以及额外的样本总计数、稀疏度参数。

方案一:改造make_classification输出复用特征结构

利用make_classification已经成熟的特征结构生成逻辑(自动区分信息/冗余/噪声特征),将其输出的连续数据转换为带0的正整数计数数据,同时控制总计数和稀疏度:

import numpy as np
from sklearn.datasets import make_classification

def make_count_classification(n_samples=1000, n_features=20, n_informative=5, 
                              n_redundant=5, n_classes=2, total_count=100, sparsity=0.3):
    # 生成带结构的连续型分类数据
    X, y = make_classification(n_samples=n_samples, n_features=n_features, 
                               n_informative=n_informative, n_redundant=n_redundant,
                               n_classes=n_classes, random_state=42)
    
    # 转换为非负值(计数不能为负)
    X_nonneg = np.abs(X)
    
    # 归一化后缩放到指定总计数并取整
    row_sums = X_nonneg.sum(axis=1, keepdims=True)
    X_normalized = X_nonneg / row_sums
    X_count = np.round(X_normalized * total_count).astype(int)
    
    # 修正取整导致的行和偏差
    diff = total_count - X_count.sum(axis=1)
    for i in range(n_samples):
        if diff[i] != 0:
            # 随机选择非0特征调整计数,确保行和等于total_count
            valid_idx = np.where(X_count[i] > 0)[0]
            if len(valid_idx) == 0:
                valid_idx = np.random.choice(n_features, 3)
            adjust_idx = np.random.choice(valid_idx, abs(diff[i]), replace=True)
            X_count[i, adjust_idx] += 1 if diff[i] > 0 else -1
    
    # 调整稀疏度:将低计数特征设为0
    flat_values = X_count.flatten()
    threshold = np.percentile(flat_values, sparsity * 100)
    X_count[X_count <= threshold] = 0
    
    # 处理调整后行和为0的样本
    zero_row_mask = X_count.sum(axis=1) == 0
    for i in np.where(zero_row_mask)[0]:
        X_count[i, np.random.choice(n_features, 3)] = np.random.randint(1, total_count//3, size=3)
    # 再次校准行和到total_count
    row_sums = X_count.sum(axis=1, keepdims=True)
    X_count = np.round(X_count / row_sums * total_count).astype(int)
    
    return X_count, y

方案说明

  • 保留了make_classification的特征结构:信息特征与标签强关联,冗余特征由信息特征线性组合生成,噪声特征无关联
  • 通过绝对值转换、归一化缩放实现从连续数据到计数数据的映射
  • 利用分位数阈值控制0值占比,精准匹配指定的稀疏度

方案二:直接生成结构化计数数据

如果希望更贴合计数数据的分布特性(比如泊松/负二项分布),可以直接构建具备结构的计数特征:

import numpy as np

def make_structured_count_data(n_samples=1000, n_features=20, n_informative=5,
                               n_redundant=5, n_classes=2, total_count=100, sparsity=0.3):
    # 生成标签
    y = np.random.randint(0, n_classes, size=n_samples)
    
    # 生成信息特征:与标签关联的泊松计数
    X_informative = []
    for cls in range(n_classes):
        # 不同类别对应不同的泊松均值,强化特征与标签的关联
        cls_means = np.random.randint(5, 20, size=n_informative)
        cls_samples = np.random.poisson(cls_means, size=(np.sum(y == cls), n_informative))
        X_informative.append(cls_samples)
    X_informative = np.vstack(X_informative)
    
    # 生成冗余特征:基于信息特征的线性组合生成计数
    X_redundant = []
    for _ in range(n_redundant):
        weights = np.random.uniform(0.2, 1.0, size=n_informative)
        linear_comb = np.dot(X_informative, weights)
        # 转换为泊松计数
        feat = np.random.poisson(linear_comb).reshape(-1, 1)
        X_redundant.append(feat)
    X_redundant = np.hstack(X_redundant)
    
    # 生成噪声特征:无关联的随机泊松计数
    n_noise = n_features - n_informative - n_redundant
    X_noise = np.random.poisson(3, size=(n_samples, n_noise))
    
    # 合并所有特征
    X = np.hstack([X_informative, X_redundant, X_noise])
    
    # 调整总计数
    row_sums = X.sum(axis=1, keepdims=True)
    X_count = np.round(X / row_sums * total_count).astype(int)
    
    # 调整稀疏度
    flat_values = X_count.flatten()
    threshold = np.percentile(flat_values, sparsity * 100)
    X_count[X_count <= threshold] = 0
    
    # 处理行和为0的样本
    zero_row_mask = X_count.sum(axis=1) == 0
    for i in np.where(zero_row_mask)[0]:
        X_count[i, np.random.choice(n_features, 3)] = np.random.randint(1, total_count//3, size=3)
    row_sums = X_count.sum(axis=1, keepdims=True)
    X_count = np.round(X_count / row_sums * total_count).astype(int)
    
    return X_count, y

方案说明

  • 信息特征直接基于标签生成,不同类别使用不同的泊松均值,确保特征与标签的关联性
  • 冗余特征由信息特征线性组合后转换为计数,符合实际场景中冗余特征的生成逻辑
  • 噪声特征采用独立泊松分布生成,模拟无意义的计数特征

内容的提问来源于stack exchange,提问作者O.rka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 00:15:09