You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中拟合多高斯曲线并提取各高斯均值?

拟合多高斯曲线并提取单峰参数(适配未知峰数量)

核心逻辑

利用高斯混合模型(GMM)自动推断数据中的峰数量,提取每个高斯成分的均值,通过相邻均值的差值估算单光电子信号值。

具体实现步骤

1. 数据预处理

先将图表数据转换为数组格式。如果是直方图数据,可生成对应样本点供GMM拟合:

import numpy as np

# 替换为你的实际数据
x = np.linspace(0, 100, 1000)
y = np.random.normal(loc=20, scale=5, size=1000) + np.random.normal(loc=40, scale=5, size=1000)

# 从直方图生成样本点(y为计数时适用)
samples = np.repeat(x, y.astype(int))

2. 自动选择GMM成分数

通过贝叶斯信息准则(BIC)遍历候选成分数,选择最优模型:

from sklearn.mixture import GaussianMixture

def get_optimal_gmm(samples, max_components=10):
    bic_scores = []
    models = []
    for n in range(1, max_components + 1):
        gmm = GaussianMixture(n_components=n, random_state=42)
        gmm.fit(samples.reshape(-1, 1))
        bic_scores.append(gmm.bic(samples.reshape(-1, 1)))
        models.append(gmm)
    # 选BIC值最小的模型
    best_idx = np.argmin(bic_scores)
    return models[best_idx]

# 获取适配数据的最优GMM模型
best_gmm = get_optimal_gmm(samples)

3. 提取峰均值并估算单光电子信号值

从GMM中提取各成分均值,计算相邻均值的差值,取中位数作为最终估算值:

# 提取并排序所有峰的均值
peak_means = np.sort(best_gmm.means_.flatten())
# 计算相邻峰的间隔
electron_value_candidates = np.diff(peak_means)
# 取中位数作为单光电子信号值(抵消个别异常间隔)
estimated_single_electron = np.median(electron_value_candidates)
print(f"估算的单光电子信号值:{estimated_single_electron:.2f}")

4. 拟合效果验证(可选)

绘制原始数据与拟合曲线,直观验证结果:

import matplotlib.pyplot as plt

def plot_fit_result(x, y, gmm):
    x_fit = np.linspace(x.min(), x.max(), 1000)
    # 计算拟合的混合高斯曲线
    fit_pdf = np.zeros_like(x_fit)
    for mean, cov, weight in zip(gmm.means_, gmm.covariances_, gmm.weights_):
        fit_pdf += weight * np.exp(-(x_fit - mean)**2 / (2 * cov)) / np.sqrt(2 * np.pi * cov)
    # 归一化拟合曲线匹配原始数据尺度
    fit_pdf = fit_pdf * (y.max() / fit_pdf.max())

    plt.figure(figsize=(10,6))
    plt.plot(x, y, label='原始数据')
    plt.plot(x_fit, fit_pdf, label='GMM拟合曲线', linestyle='--')
    # 标记每个峰的均值位置
    for idx, mean in enumerate(gmm.means_):
        plt.axvline(x=mean, color='r', linestyle=':', label='峰均值' if idx == 0 else "")
    plt.legend()
    plt.xlabel('信号强度')
    plt.ylabel('计数')
    plt.show()

# 调用绘图函数
plot_fit_result(x, y, best_gmm)

注意事项

  • 若原始数据是带权重的直方图,可直接使用GaussianMixture.fit()的sample_weight参数传入权重,无需生成样本点。
  • max_components需根据实际数据调整,比如光电子峰数量通常不会过多,可设为5-10。
  • 若峰重叠严重导致BIC低估成分数,可结合光电子峰间隔大致固定的先验知识,手动微调成分数。

内容的提问来源于stack exchange,提问作者junoswrld

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 03:55:29