You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效拟合Length-Count格式大型数据集的分布?

高效分析带频次列的数据集分布方法

核心思路:无需展开数据集,直接将Count列作为权重参与所有分布分析步骤,彻底避免大规模数据展开带来的内存占用和时间损耗。

1. Scipy分布拟合直接传入权重

Scipy的绝大多数分布拟合函数都支持weights参数,直接传入Count列即可完成拟合,无需生成重复的一维列表:

连续分布拟合示例(以正态分布为例)

import numpy as np
import pandas as pd
from scipy import stats

# 加载你的数据集
df = pd.DataFrame({'length': [5, 10], 'Count': [12, 4]})
x = df['length'].values
weights = df['Count'].values

# 拟合正态分布,传入weights参数
loc, scale = stats.norm.fit(x, weights=weights)
print(f"拟合得到正态分布参数:均值={loc:.2f},标准差={scale:.2f}")

离散分布拟合示例(以泊松分布为例)

lambda_hat = stats.poisson.fit(x, weights=weights)[0]
print(f"拟合得到泊松分布参数λ={lambda_hat:.2f}")

2. 计算加权统计量

先计算基础分布特征时,直接用加权版本的统计方法,效率远超展开数据:

  • 加权均值:np.average(df['length'], weights=df['Count'])
  • 加权方差:
mean = np.average(df['length'], weights=df['Count'])
var = np.average((df['length'] - mean)**2, weights=df['Count'])
  • 加权分位数:
from scipy.stats.mstats import mquantiles

# 计算四分位数
weighted_quarts = mquantiles(df['length'], prob=[0.25, 0.5, 0.75], weights=df['Count'])

3. 加权可视化

绘制分布图表时,同样通过权重参数实现,无需展开数据:

加权直方图

import matplotlib.pyplot as plt

plt.hist(df['length'], weights=df['Count'], bins='auto', edgecolor='black')
plt.xlabel('Length')
plt.ylabel('Frequency')
plt.title('Weighted Histogram of Length')
plt.show()

加权核密度估计(KDE)

from statsmodels.nonparametric.kernel_density import KDEMultivariate

# 初始化加权KDE模型
kde = KDEMultivariate(df['length'], var_type='c', weights=df['Count'])
# 生成绘图用的x轴网格
x_grid = np.linspace(df['length'].min(), df['length'].max(), 1000)
# 计算密度值
density = kde.pdf(x_grid)

plt.plot(x_grid, density, linewidth=2)
plt.xlabel('Length')
plt.ylabel('Density')
plt.title('Weighted KDE of Length')
plt.show()

内容的提问来源于stack exchange,提问作者Disha Verma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 18:12:59