如何高效拟合Length-Count格式大型数据集的分布?
高效分析带频次列的数据集分布方法
核心思路:无需展开数据集,直接将Count列作为权重参与所有分布分析步骤,彻底避免大规模数据展开带来的内存占用和时间损耗。
1. Scipy分布拟合直接传入权重
Scipy的绝大多数分布拟合函数都支持weights参数,直接传入Count列即可完成拟合,无需生成重复的一维列表:
连续分布拟合示例(以正态分布为例)
import numpy as np import pandas as pd from scipy import stats # 加载你的数据集 df = pd.DataFrame({'length': [5, 10], 'Count': [12, 4]}) x = df['length'].values weights = df['Count'].values # 拟合正态分布,传入weights参数 loc, scale = stats.norm.fit(x, weights=weights) print(f"拟合得到正态分布参数:均值={loc:.2f},标准差={scale:.2f}")
离散分布拟合示例(以泊松分布为例)
lambda_hat = stats.poisson.fit(x, weights=weights)[0] print(f"拟合得到泊松分布参数λ={lambda_hat:.2f}")
2. 计算加权统计量
先计算基础分布特征时,直接用加权版本的统计方法,效率远超展开数据:
- 加权均值:
np.average(df['length'], weights=df['Count']) - 加权方差:
mean = np.average(df['length'], weights=df['Count']) var = np.average((df['length'] - mean)**2, weights=df['Count'])
- 加权分位数:
from scipy.stats.mstats import mquantiles # 计算四分位数 weighted_quarts = mquantiles(df['length'], prob=[0.25, 0.5, 0.75], weights=df['Count'])
3. 加权可视化
绘制分布图表时,同样通过权重参数实现,无需展开数据:
加权直方图
import matplotlib.pyplot as plt plt.hist(df['length'], weights=df['Count'], bins='auto', edgecolor='black') plt.xlabel('Length') plt.ylabel('Frequency') plt.title('Weighted Histogram of Length') plt.show()
加权核密度估计(KDE)
from statsmodels.nonparametric.kernel_density import KDEMultivariate # 初始化加权KDE模型 kde = KDEMultivariate(df['length'], var_type='c', weights=df['Count']) # 生成绘图用的x轴网格 x_grid = np.linspace(df['length'].min(), df['length'].max(), 1000) # 计算密度值 density = kde.pdf(x_grid) plt.plot(x_grid, density, linewidth=2) plt.xlabel('Length') plt.ylabel('Density') plt.title('Weighted KDE of Length') plt.show()
内容的提问来源于stack exchange,提问作者Disha Verma
相关产品推荐
相关产品推荐

