You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用mclust包Mclust函数时出现方差/σ无穷值异常结果求助

Fixing Infinite Variance Issues with Mclust on Discrete Rating Data

Hey there! I’ve run into this exact problem before when using Gaussian-based mixed models like Mclust on integer rating data (like your 1-7 scale columns). Let’s walk through why this is happening and how to fix it:

Why You’re Seeing Infinite Variance

The diabetes vignette uses continuous physiological data (e.g., blood glucose levels) where Gaussian assumptions hold well. But your data is discrete, ordered rating data—and worse, some rows/columns have identical values (like row 3 in your sample, where every entry is 7). When Mclust tries to fit a Gaussian model to a subset of data where a variable has zero variance (all values the same), it can’t compute a valid covariance matrix (since division by zero leads to infinite values).

Step-by-Step Fixes

1. First: Hunt for Zero-Variance Variables or Uniform Subgroups

Start by checking which columns have no variation, or if any potential clusters would have uniform values:

# Check variance of each column
sapply(example, var)
# Check for rows with all identical values (like row 3 in your sample)
row_all_same <- apply(example, 1, function(x) length(unique(x)) == 1)
table(row_all_same)
  • If you find columns with zero variance, drop them—they add no information and break the model.
  • If you have rows with all identical values, consider removing them (if they’re outliers) or grouping them manually before fitting the model.

2. Switch to a Model Built for Discrete Data

Mclust is designed for continuous data. For your ordered rating data, Latent Class Analysis (LCA) is a better fit. The poLCA package is made exactly for this:

library(poLCA)
# Convert all columns to factors (required for poLCA)
example_factors <- as.data.frame(lapply(example, as.factor))
# Define the LCA model (all variables as indicators of latent classes)
lca_formula <- cbind(HI01_01, HI01_02, HI01_03, HI01_07, HI01_08, HI01_09, HI01_10, HI01_12) ~ 1
# Fit a 3-class model (adjust nclass to test different numbers of classes)
lca_model <- poLCA(lca_formula, data = example_factors, nclass = 3)
# View results
summary(lca_model)

3. If You Must Use Mclust: Constrain the Variance Structure

If you need to stick with mclust, avoid the default full covariance matrix ("VVV"), which is too flexible for discrete data. Instead, use a constrained variance structure:

library(mclust)
# Try spherical variance (all variables have same variance)
model_eii <- Mclust(example, modelNames = "EII", G = 2:5)
# Or diagonal variance (each variable has its own variance, no covariance)
model_vii <- Mclust(example, modelNames = "VII", G = 2:5)
# Check which model fits best
summary(model_eii)
summary(model_vii)

Constrained models require fewer parameters, so they’re more stable with discrete or small datasets.

4. Preprocess Your Data

Standardizing your rating data can help stabilize parameter estimates:

# Standardize each column to mean 0, variance 1
example_scaled <- scale(example)
# Fit Mclust on scaled data
model_scaled <- Mclust(example_scaled, G = 2:5)

Just remember to interpret results relative to the scaled values, or reverse-transform if needed.

Quick Check: Sample Size per Cluster

If your model is trying to fit a cluster with only 1-2 observations, variance estimates will be unstable. Use model$classification to check cluster sizes after fitting, and adjust the number of classes (G) if needed.

内容的提问来源于stack exchange,提问作者Saja01

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:04:54