You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何绘制贝叶斯高斯混合模型的平滑连续概率密度函数(PDF)

问题原因

  • 你的输入数据集是2维特征,原代码仅对少量训练样本点计算PDF,且自行生成的1维采样序列x = np.linspace(0, 1, num=20)和数据实际的特征取值范围(0~9区间)完全不匹配,强行对齐绘图只能得到离散结果。
  • 要得到平滑的PDF,需要生成覆盖数据全范围的足够密集的采样点,再用训练好的模型计算这些采样点的PDF值再绘图。

修改后可运行代码

import numpy as np
import matplotlib.pyplot as plt
from sklearn.mixture import BayesianGaussianMixture

# 直接使用你提供的数据集,也可以替换为读本地文件的逻辑
data = np.array([
    [6.11507621, 6.2285484 ],
    [5.61154419, 7.4166868 ],
    [5.3638034,  8.64581576],
    [8.58030274, 6.01384676],
    [2.06883754, 8.5662325 ],
    [7.772149,   2.29177372],
    [0.66223423, 0.01642353],
    [7.42461573, 5.46288677],
    [0.82355307, 3.60322705],
    [1.12966405, 9.54888118],
    [4.34716189, 3.63203485],
    [7.95368286, 5.74659859],
    [3.21564946, 3.67576324],
    [6.48021187, 7.35190659],
    [3.02668358, 4.41981514],
    [0.01745485, 7.49153586],
    [1.08490595, 0.91004064],
    [1.89995405, 0.38728879],
    [4.40549506, 2.48715052],
    [4.52857064, 1.24935027]
])
X_train = data

# 模型训练逻辑和原代码一致
bgmm = BayesianGaussianMixture(
    n_components=15,
    random_state=7,
    max_iter=5000,
    n_init=10,
    weight_concentration_prior_type="dirichlet_distribution"
)
bgmm.fit(X_train)

# --------------------------
# 方案1:绘制2维联合分布平滑等高线图
# --------------------------
# 生成密集网格采样点,采样范围匹配数据实际取值
x1_min, x1_max = X_train[:, 0].min() - 1, X_train[:, 0].max() + 1
x2_min, x2_max = X_train[:, 1].min() - 1, X_train[:, 1].max() + 1
xx1, xx2 = np.meshgrid(
    np.linspace(x1_min, x1_max, 100),
    np.linspace(x2_min, x2_max, 100)
)
grid_samples = np.c_[xx1.ravel(), xx2.ravel()]
# 计算所有网格点的PDF
logprob = bgmm.score_samples(grid_samples)
pdf = np.exp(logprob).reshape(xx1.shape)

# 绘制等高线+原始样本点
plt.figure(figsize=(8, 6))
plt.contourf(xx1, xx2, pdf, cmap='Blues')
plt.colorbar(label='PDF值')
plt.scatter(X_train[:, 0], X_train[:, 1], c='red', s=20, label='原始样本点')
plt.xlabel('特征1')
plt.ylabel('特征2')
plt.legend()
plt.title('贝叶斯高斯混合模型2维联合PDF')
plt.show()

# --------------------------
# 方案2:绘制单特征的1维边缘平滑PDF曲线
# --------------------------
def cal_1d_pdf(samples, weights, means, covs, dim):
    """计算指定维度的边缘PDF"""
    pdf = np.zeros_like(samples)
    for w, mu, cov in zip(weights, means, covs):
        mu_dim = mu[dim]
        cov_dim = cov[dim, dim]
        pdf += w * (1/(np.sqrt(2*np.pi*cov_dim))) * np.exp(-(samples - mu_dim)**2/(2*cov_dim))
    return pdf

plt.figure(figsize=(12, 5))
# 特征1的边缘PDF
plt.subplot(1, 2, 1)
x1_samples = np.linspace(x1_min, x1_max, 200)
x1_pdf = cal_1d_pdf(x1_samples, bgmm.weights_, bgmm.means_, bgmm.covariances_, 0)
plt.plot(x1_samples, x1_pdf, 'k-', linewidth=2)
plt.xlabel('特征1')
plt.ylabel('PDF值')
plt.title('特征1边缘分布PDF')

# 特征2的边缘PDF
plt.subplot(1, 2, 2)
x2_samples = np.linspace(x2_min, x2_max, 200)
x2_pdf = cal_1d_pdf(x2_samples, bgmm.weights_, bgmm.means_, bgmm.covariances_, 1)
plt.plot(x2_samples, x2_pdf, 'k-', linewidth=2)
plt.xlabel('特征2')
plt.ylabel('PDF值')
plt.title('特征2边缘分布PDF')
plt.tight_layout()
plt.show()

关键修改说明

  • 按数据集两个特征的实际取值范围生成了密集采样点,每个维度采样100~200个点,足够的采样密度保证了绘制的PDF是平滑的
  • 提供两种可视化方案:需要查看两个特征的联合分布选择方案1,需要查看单个特征的分布曲线选择方案2
  • 原代码中X_train = np.reshape(data, (10*data.shape[0],2))是将原始20条数据重复10次,如果你是为了做数据增强可以保留该逻辑,否则直接使用原始数据训练即可。

内容的提问来源于stack exchange,提问作者user1393214

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 14:36:07