如何基于类别生成符合原分布的连续型price模拟数据
代码实现方案
核心逻辑是先按类别拟合原始价格的分布特征,再基于拟合的分布为每个类别生成对应数量的随机值,有两种常用实现路径:
方案1:非参数核密度拟合(无需假设分布类型,最贴合原始分布)
不需要提前假设价格服从什么分布,完全基于原始数据的分布特征拟合,适用所有场景。
步骤1:导入依赖库
import pandas as pd import numpy as np from scipy.stats import gaussian_kde
步骤2:按类别拟合分布
# 存储每个类别的分布参数 category_dist_map = {} for cate, group_df in original_df.groupby("category name"): # 过滤空值 raw_prices = group_df["price"].dropna().values # 拟合核密度模型 kde_model = gaussian_kde(raw_prices) # 记录原始价格的上下限,避免生成异常值 price_min = raw_prices.min() price_max = raw_prices.max() category_dist_map[cate] = { "kde": kde_model, "min": price_min, "max": price_max, "sample_num": len(group_df) }
步骤3:生成新的价格列
如果不需要生成的价格完全唯一,直接用下面的代码:
new_price_list = [] for cate, dist_info in category_dist_map.items(): # 生成和该类别样本量一致的随机值 random_prices = dist_info["kde"].resample(dist_info["sample_num"])[0] # 裁剪到原始价格范围内 random_prices = np.clip(random_prices, dist_info["min"], dist_info["max"]) # 保留2位小数(可根据业务调整精度) random_prices = np.round(random_prices, 2) new_price_list.extend([(cate, p) for p in random_prices]) # 合并到原始DataFrame new_price_df = pd.DataFrame(new_price_list, columns=["category name", "new_price"]) original_df = original_df.reset_index()\ .merge(new_price_df, on="category name", how="left")\ .set_index("index")
如果要求生成的所有价格完全不重复,把生成随机值的部分替换为以下逻辑:
new_price_list = [] for cate, dist_info in category_dist_map.items(): sample_num = dist_info["sample_num"] unique_prices = set() res = [] while len(res) < sample_num: p = dist_info["kde"].resample(1)[0][0] p = round(np.clip(p, dist_info["min"], dist_info["max"]), 2) if p not in unique_prices: unique_prices.add(p) res.append(p) new_price_list.extend([(cate, p) for p in res])
方案2:参数分布拟合(性能更高,适合价格符合常见分布的场景)
如果你的价格是右偏的电商类价格,绝大多数符合对数正态分布,用这个方案速度比核密度快3~5倍:
from scipy.stats import lognorm # 拟合分布 category_dist_map = {} for cate, group_df in original_df.groupby("category name"): raw_prices = group_df["price"].dropna().values # 拟合对数正态分布参数 s, loc, scale = lognorm.fit(raw_prices) category_dist_map[cate] = { "s": s, "loc": loc, "scale": scale, "min": raw_prices.min(), "max": raw_prices.max(), "sample_num": len(group_df) } # 生成随机值 new_price_list = [] for cate, dist_info in category_dist_map.items(): random_prices = lognorm.rvs( s=dist_info["s"], loc=dist_info["loc"], scale=dist_info["scale"], size=dist_info["sample_num"] ) random_prices = np.round(np.clip(random_prices, dist_info["min"], dist_info["max"]), 2) new_price_list.extend([(cate, p) for p in random_prices]) # 合并逻辑和方案1一致 new_price_df = pd.DataFrame(new_price_list, columns=["category name", "new_price"]) original_df = original_df.reset_index()\ .merge(new_price_df, on="category name", how="left")\ .set_index("index")
注意事项
- 如果某个类别的样本量少于30,核密度拟合效果会变差,建议直接从该类别的原始价格中无放回采样即可
- 小数位精度可以根据实际业务调整,比如整数价格就把round的参数改为0
内容的提问来源于stack exchange,提问作者kurumi
相关产品推荐
相关产品推荐

