含nominal、ordinal与metric变量的数据集聚类分析方法咨询
Great question! K-modes is totally a reasonable starting point when dealing with mixed-scale data—it’s built to handle nominal variables by leaning into mode-based similarity, which makes sense for categorical data. But depending on your dataset’s specifics and what you’re trying to get out of your clustering, there are some stronger alternatives worth exploring:
1. K-Prototypes Clustering
This is essentially the "upgrade" to k-modes for mixed data—it combines the best parts of k-means (for metric variables) and k-modes (for nominal/ordinal). It uses Euclidean distance for numerical features and Hamming distance for categorical ones, with a tunable weight to balance their impact.
- Why it’s better than k-modes: It doesn’t throw away the valuable numerical information in your metric variables, which k-modes ignores entirely.
- Pro tip: Play around with the weight parameter to prioritize feature types that matter most for your clustering goal. For example, if your metric variables are more predictive of groupings, crank up their weight.
2. Hierarchical Clustering with Custom Distances
Hierarchical clustering lets you define a distance metric that works for all your variable types, and Gower’s distance is the go-to here. It normalizes each feature category:
- For metric variables: Uses z-scores to standardize values
- For nominal/ordinal: Uses matching scores (1 if values are the same, 0 otherwise, adjusted for ordinal ranks)
Then it combines these into a single distance value between 0 and 1. - Pros: You don’t have to guess the number of clusters upfront—the dendrogram will help you visualize how groups merge.
- Cons: It can get slow if you have a really large dataset, since it computes pairwise distances for every data point.
3. Model-Based Clustering (Extended Mixture Models)
If you’re comfortable with probabilistic approaches, you can use mixture models tailored to mixed data. For example:
- Metric variables follow a Gaussian distribution
- Nominal variables follow a multinomial distribution
- Ordinal variables use ordinal logistic distributions
These models assign each data point a probability of belonging to each cluster, which can help with uncertainty. - Pros: Gives you statistical rigor and handles ambiguous cases well.
- Cons: It’s more computationally heavy, and you have to make sure your distribution assumptions match your data (which isn’t always straightforward).
4. Fuzzy Mixed Data Clustering
Fuzzy versions of k-prototypes or k-modes let data points belong to multiple clusters with different membership degrees. This is perfect if your data has overlapping groups where a single hard assignment doesn’t make sense.
- Pros: Captures the messiness of real-world data better than hard clustering methods.
- Cons: Interpreting results is trickier—you’ll need to set thresholds for membership to define "true" clusters.
My Quick Take
If your dataset has both metric and categorical variables, k-prototypes should be your first stop instead of k-modes. It’s more flexible and uses all your data. If you’re unsure about the number of clusters or want to explore hierarchical relationships, go with hierarchical clustering using Gower’s distance.
Before settling on a method, do these checks:
- Test a couple of approaches and compare cluster quality with metrics like the silhouette score (adapted for mixed data) or external validation if you have ground truth labels.
- Always normalize your metric variables first (e.g.,
z-score) so they don’t overpower the categorical features in distance calculations.
内容的提问来源于stack exchange,提问作者Papea

