基于轮廓聚类准则评估TDA Mapper聚类输出的方法问询
Great question—TDA Mapper’s overlapping cluster structure does break the assumptions of traditional metrics like silhouette scores, which rely on mutually exclusive clusters. Here are practical, tailored approaches to evaluate its output, plus ways to compare it against your hierarchical and mclust results:
1. Modified Silhouette Scores for Overlapping Assignments
Traditional silhouette scores calculate a single value per sample based on one cluster membership. For Mapper, since a sample can belong to multiple clusters, adjust the metric like this:
- For each sample
iand every clusterCit’s part of, compute a cluster-specific silhouette score:
wheres(i,C) = (b(i,C) - a(i,C)) / max(a(i,C), b(i,C))a(i,C)is the average distance fromito other samples in clusterC, andb(i,C)is the average distance fromito the nearest cluster thatidoes not belong to. - For each sample, take the average (or maximum) of its
s(i,C)values across all clusters it’s in. Then compute the overall average across all samples to get a single score for your Mapper run. - This adjusted score lets you directly compare Mapper against your hierarchical clustering and mclust results, which use standard silhouette scores.
2. Cluster Stability & Robustness Checks
Mapper results are highly sensitive to parameters (filter function, resolution, overlap percentage). Evaluating stability is a strong way to validate its utility:
- Run Mapper multiple times with small tweaks to parameters (e.g., slight shifts in resolution, different filter functions if applicable).
- Use metrics like Jaccard similarity to compare cluster overlap across runs: measure how often pairs of samples co-occur in the same cluster across different parameter settings.
- Compare this stability to your hierarchical clustering (e.g., how consistent cluster assignments are when changing linkage methods) and mclust (e.g., how stable model selection via BIC is across subsampled data). More stable results often indicate better alignment with underlying data structure.
3. Domain-Specific Validation (If You Have Labels)
If your dataset includes ground-truth labels (e.g., known classes), leverage them to evaluate Mapper’s overlapping clusters:
- Cluster Purity for Overlaps: For each cluster, calculate the proportion of samples belonging to the most frequent true label. Then, for each sample, average the purity scores of all clusters it’s part of. Higher average purity means Mapper’s clusters align well with known groups.
- Coverage Accuracy: Count how many samples are part of at least one cluster where their true label is the majority class. The percentage of such samples gives a straightforward measure of how well Mapper captures meaningful groups, even with overlaps.
- Contrast this with the standard purity or accuracy scores you’d use for hierarchical clustering and mclust to see which method better matches your domain knowledge.
4. Topological Validity Metrics
Since Mapper is rooted in topological data analysis, use topology-aligned metrics to assess its output:
- Graph Modularity: Treat Mapper’s output as a graph (nodes = clusters, edges = overlapping samples). Compute modularity to measure how well the graph’s communities (connected clusters) group similar samples. Higher modularity indicates more meaningful topological structure.
- Density Consistency: For each Mapper cluster, calculate the average pairwise distance between samples (lower = denser). Check if denser clusters correspond to intuitively meaningful groups in your data, and compare this to the cluster density patterns from hierarchical clustering/mclust.
- You can also validate if Mapper’s graph structure reflects known topological features of your data (e.g., if your data has a circular shape, does Mapper produce a cyclic graph?).
5. Cross-Method Comparison Metrics
To directly compare Mapper against hierarchical clustering and mclust (which produce non-overlapping clusters):
- Overlapping Mutual Information (OMI): A variant of mutual information designed for overlapping clusters. It quantifies the similarity between Mapper’s overlapping assignments and the non-overlapping assignments from your other methods. Higher OMI means better alignment.
- Co-Occurrence Agreement: For every pair of samples, count how often they’re in the same cluster in Mapper, versus how often they’re in the same cluster in hierarchical clustering/mclust. Compute the correlation between these two counts to measure consistency across methods.
Content of the question comes from Stack Exchange, asked by Katherina

