You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于类别生成符合原分布的连续型price模拟数据

代码实现方案

核心逻辑是先按类别拟合原始价格的分布特征,再基于拟合的分布为每个类别生成对应数量的随机值,有两种常用实现路径:

方案1:非参数核密度拟合(无需假设分布类型,最贴合原始分布)

不需要提前假设价格服从什么分布,完全基于原始数据的分布特征拟合,适用所有场景。

步骤1:导入依赖库

import pandas as pd
import numpy as np
from scipy.stats import gaussian_kde

步骤2:按类别拟合分布

# 存储每个类别的分布参数
category_dist_map = {}

for cate, group_df in original_df.groupby("category name"):
    # 过滤空值
    raw_prices = group_df["price"].dropna().values
    # 拟合核密度模型
    kde_model = gaussian_kde(raw_prices)
    # 记录原始价格的上下限,避免生成异常值
    price_min = raw_prices.min()
    price_max = raw_prices.max()
    category_dist_map[cate] = {
        "kde": kde_model,
        "min": price_min,
        "max": price_max,
        "sample_num": len(group_df)
    }

步骤3:生成新的价格列

如果不需要生成的价格完全唯一,直接用下面的代码:

new_price_list = []
for cate, dist_info in category_dist_map.items():
    # 生成和该类别样本量一致的随机值
    random_prices = dist_info["kde"].resample(dist_info["sample_num"])[0]
    # 裁剪到原始价格范围内
    random_prices = np.clip(random_prices, dist_info["min"], dist_info["max"])
    # 保留2位小数(可根据业务调整精度)
    random_prices = np.round(random_prices, 2)
    new_price_list.extend([(cate, p) for p in random_prices])

# 合并到原始DataFrame
new_price_df = pd.DataFrame(new_price_list, columns=["category name", "new_price"])
original_df = original_df.reset_index()\
    .merge(new_price_df, on="category name", how="left")\
    .set_index("index")

如果要求生成的所有价格完全不重复,把生成随机值的部分替换为以下逻辑:

new_price_list = []
for cate, dist_info in category_dist_map.items():
    sample_num = dist_info["sample_num"]
    unique_prices = set()
    res = []
    while len(res) < sample_num:
        p = dist_info["kde"].resample(1)[0][0]
        p = round(np.clip(p, dist_info["min"], dist_info["max"]), 2)
        if p not in unique_prices:
            unique_prices.add(p)
            res.append(p)
    new_price_list.extend([(cate, p) for p in res])

方案2:参数分布拟合(性能更高,适合价格符合常见分布的场景)

如果你的价格是右偏的电商类价格,绝大多数符合对数正态分布,用这个方案速度比核密度快3~5倍:

from scipy.stats import lognorm

# 拟合分布
category_dist_map = {}
for cate, group_df in original_df.groupby("category name"):
    raw_prices = group_df["price"].dropna().values
    # 拟合对数正态分布参数
    s, loc, scale = lognorm.fit(raw_prices)
    category_dist_map[cate] = {
        "s": s, "loc": loc, "scale": scale,
        "min": raw_prices.min(), "max": raw_prices.max(),
        "sample_num": len(group_df)
    }

# 生成随机值
new_price_list = []
for cate, dist_info in category_dist_map.items():
    random_prices = lognorm.rvs(
        s=dist_info["s"], loc=dist_info["loc"], scale=dist_info["scale"],
        size=dist_info["sample_num"]
    )
    random_prices = np.round(np.clip(random_prices, dist_info["min"], dist_info["max"]), 2)
    new_price_list.extend([(cate, p) for p in random_prices])

# 合并逻辑和方案1一致
new_price_df = pd.DataFrame(new_price_list, columns=["category name", "new_price"])
original_df = original_df.reset_index()\
    .merge(new_price_df, on="category name", how="left")\
    .set_index("index")

注意事项

  • 如果某个类别的样本量少于30,核密度拟合效果会变差,建议直接从该类别的原始价格中无放回采样即可
  • 小数位精度可以根据实际业务调整,比如整数价格就把round的参数改为0

内容的提问来源于stack exchange,提问作者kurumi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 08:24:08