如何绘制贝叶斯高斯混合模型的平滑连续概率密度函数(PDF)
问题原因
- 你的输入数据集是2维特征,原代码仅对少量训练样本点计算PDF,且自行生成的1维采样序列
x = np.linspace(0, 1, num=20)和数据实际的特征取值范围(0~9区间)完全不匹配,强行对齐绘图只能得到离散结果。 - 要得到平滑的PDF,需要生成覆盖数据全范围的足够密集的采样点,再用训练好的模型计算这些采样点的PDF值再绘图。
修改后可运行代码
import numpy as np import matplotlib.pyplot as plt from sklearn.mixture import BayesianGaussianMixture # 直接使用你提供的数据集,也可以替换为读本地文件的逻辑 data = np.array([ [6.11507621, 6.2285484 ], [5.61154419, 7.4166868 ], [5.3638034, 8.64581576], [8.58030274, 6.01384676], [2.06883754, 8.5662325 ], [7.772149, 2.29177372], [0.66223423, 0.01642353], [7.42461573, 5.46288677], [0.82355307, 3.60322705], [1.12966405, 9.54888118], [4.34716189, 3.63203485], [7.95368286, 5.74659859], [3.21564946, 3.67576324], [6.48021187, 7.35190659], [3.02668358, 4.41981514], [0.01745485, 7.49153586], [1.08490595, 0.91004064], [1.89995405, 0.38728879], [4.40549506, 2.48715052], [4.52857064, 1.24935027] ]) X_train = data # 模型训练逻辑和原代码一致 bgmm = BayesianGaussianMixture( n_components=15, random_state=7, max_iter=5000, n_init=10, weight_concentration_prior_type="dirichlet_distribution" ) bgmm.fit(X_train) # -------------------------- # 方案1:绘制2维联合分布平滑等高线图 # -------------------------- # 生成密集网格采样点,采样范围匹配数据实际取值 x1_min, x1_max = X_train[:, 0].min() - 1, X_train[:, 0].max() + 1 x2_min, x2_max = X_train[:, 1].min() - 1, X_train[:, 1].max() + 1 xx1, xx2 = np.meshgrid( np.linspace(x1_min, x1_max, 100), np.linspace(x2_min, x2_max, 100) ) grid_samples = np.c_[xx1.ravel(), xx2.ravel()] # 计算所有网格点的PDF logprob = bgmm.score_samples(grid_samples) pdf = np.exp(logprob).reshape(xx1.shape) # 绘制等高线+原始样本点 plt.figure(figsize=(8, 6)) plt.contourf(xx1, xx2, pdf, cmap='Blues') plt.colorbar(label='PDF值') plt.scatter(X_train[:, 0], X_train[:, 1], c='red', s=20, label='原始样本点') plt.xlabel('特征1') plt.ylabel('特征2') plt.legend() plt.title('贝叶斯高斯混合模型2维联合PDF') plt.show() # -------------------------- # 方案2:绘制单特征的1维边缘平滑PDF曲线 # -------------------------- def cal_1d_pdf(samples, weights, means, covs, dim): """计算指定维度的边缘PDF""" pdf = np.zeros_like(samples) for w, mu, cov in zip(weights, means, covs): mu_dim = mu[dim] cov_dim = cov[dim, dim] pdf += w * (1/(np.sqrt(2*np.pi*cov_dim))) * np.exp(-(samples - mu_dim)**2/(2*cov_dim)) return pdf plt.figure(figsize=(12, 5)) # 特征1的边缘PDF plt.subplot(1, 2, 1) x1_samples = np.linspace(x1_min, x1_max, 200) x1_pdf = cal_1d_pdf(x1_samples, bgmm.weights_, bgmm.means_, bgmm.covariances_, 0) plt.plot(x1_samples, x1_pdf, 'k-', linewidth=2) plt.xlabel('特征1') plt.ylabel('PDF值') plt.title('特征1边缘分布PDF') # 特征2的边缘PDF plt.subplot(1, 2, 2) x2_samples = np.linspace(x2_min, x2_max, 200) x2_pdf = cal_1d_pdf(x2_samples, bgmm.weights_, bgmm.means_, bgmm.covariances_, 1) plt.plot(x2_samples, x2_pdf, 'k-', linewidth=2) plt.xlabel('特征2') plt.ylabel('PDF值') plt.title('特征2边缘分布PDF') plt.tight_layout() plt.show()
关键修改说明
- 按数据集两个特征的实际取值范围生成了密集采样点,每个维度采样100~200个点,足够的采样密度保证了绘制的PDF是平滑的
- 提供两种可视化方案:需要查看两个特征的联合分布选择方案1,需要查看单个特征的分布曲线选择方案2
- 原代码中
X_train = np.reshape(data, (10*data.shape[0],2))是将原始20条数据重复10次,如果你是为了做数据增强可以保留该逻辑,否则直接使用原始数据训练即可。
内容的提问来源于stack exchange,提问作者user1393214
相关产品推荐
相关产品推荐

