You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python计算指定点附近数据分布的质量/密度?

计算指定区间内的数据分布质量与密度

数据集

ll = [25.553885868617463,
 1.4285714285714288,
 5.0,
 14.142857142857142,
 2.714285430908202,
 -4.428571428571429,
 3.428571428571429,
 2.8571428571428568,
 5.857142857142858,
 -2.0,
 8.571428571428573,
 1.4285714285714288,
 1.857142333984374,
 21.714285714285715,
 3.1428571428571423,
 2.428571428571427,
 -0.2857142857142856,
 -3.0,
 4.142857687813894,
 0.7142857142857135,
 0.714285507202149,
 -1.9999995858328674,
 4.857142464773997,
 2.8571428571428577,
 -6.714285714285714,
 3.57142848423549,
 15.999999901907785,
 0.14285714285714413,
 -3.0,
 0.5687830243791847,
 7.857142900739401,
 -3.0,
 9.0,
 2.428571428571427,
 2.0000001634870266,
 0.7999999999999998,
 -5.7142857142857135,
 3.1428571428571423,
 0.14285714285714235,
 22.5,
 18.571428527832033,
 2.7142857142857135,
 0,
 3.1428571428571423,
 13.142857142857146,
 10.428571428571427,
 30.71428684779576,
 0,
 0.2857140350341787,
 3.571428571428571,
 2.0,
 24.428570175170897,
 2.428571428571429,
 -0.3333333333333339,
 4.2857142857142865,
 -8.000000216166178,
 15.57142857142857,
 2.2857142857142856,
 8.71428565979004,
 0.8571428571428577,
 2.1428570447649276,
 1.0,
 5.000000991821288,
 4.714285714285715,
 6.0,
 2.8571428571428577,
 1.6666666666666679,
 1.9987989153180798,
 12.714285714285715,
 9.85714340209961,
 7.71428658621652,
 -5.857142857142858,
 15.857142857142858,
 4.428571428571429,
 0.5676193237304688,
 1.2857142857142847,
 0.14285705566406248,
 3.428570938110351,
 5.142857142857142,
 -1.2857142857142856,
 -1.0,
 11.714285714285715,
 -0.7142857142857144,
 0.714285888671875,
 -1.0,
 9.428571428571429,
 4.428571428571429,
 -2.428571428571429,
 -20.571428571428573,
 4.0,
 1.1428571428571432,
 2.2857142857142847,
 19.0,
 15.142857142857142,
 5.571428451538086,
 7.428571428571427,
 1.0,
 4.285714481898715,
 3.7142853546142582,
 -3.7142854309082036]

绘制基础直方图

使用以下代码生成数据的直方图:

import pandas as pd
import plotly.express as px

px.histogram(pd.DataFrame(ll, columns=['val']), x='val', nbins=100)

直方图

需求说明

我们需要针对任意指定的centre(中心值)和thr(区间半宽),计算落在[centre-thr, centre+thr]区间内的数据分布质量(该区间内数据点占总数据的比例),以及分布密度(单位区间长度内的质量占比)。示例代码中两条红线标记了目标区间:

thr = 0.5
centre = 0
fig = px.histogram(pd.DataFrame(ll, columns=['val']), x='val', nbins=100)
fig.add_vline(x=centre+thr, line_color='red')
fig.add_vline(x=centre-thr, line_color='red')
fig.show()

带红线的直方图

计算方法与代码实现

核心逻辑

  • 分布质量:区间内数据点数量 ÷ 总数据点数量,代表该区间内数据的占比
  • 分布密度:分布质量 ÷ 区间总长度(即2*thr),代表单位长度内的数据占比

完整代码

import pandas as pd

# 将列表转换为DataFrame
df = pd.DataFrame(ll, columns=['val'])

# 定义目标区间参数
thr = 0.5
centre = 0

# 计算区间上下限
lower_bound = centre - thr
upper_bound = centre + thr

# 统计区间内的样本数量
in_range_count = df[(df['val'] >= lower_bound) & (df['val'] <= upper_bound)].shape[0]
total_count = df.shape[0]

# 计算分布质量(占比)
mass = in_range_count / total_count
# 计算分布密度
density = mass / (2 * thr)

print(f"区间[{lower_bound}, {upper_bound}]内的分布质量:{mass:.4f}")
print(f"区间[{lower_bound}, {upper_bound}]内的分布密度:{density:.4f}")

结果说明

运行上述代码后,会输出指定区间内的数据占比和单位长度占比。你可以直接修改thr和centre的值,快速计算任意目标区间的分布指标。

内容的提问来源于stack exchange,提问作者quant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 09:55:59