You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TraMineR聚类得分获取方法咨询(R3.4.4环境)

解答:获取个体聚类得分并分析极端值

Hi Silvia,

Absolutely! You can calculate scores that measure how "extreme" each individual is within its cluster—perfect for your analysis. Since you're working with sequence clustering using TraMineR and hierarchical clustering (ward method via agnes), we’ll use your original sequence distance matrix (dist.om2) to compute meaningful scores tied directly to your clustering logic.

核心思路

For your ward-based hierarchical clustering, the most intuitive "clustering score" is a measure of how far an individual is from its cluster’s center, or how dissimilar it is to other members of the same cluster. Larger values mean the individual is more extreme (less aligned with the cluster’s core).

具体实现步骤

Let’s integrate this into your existing workflow seamlessly:

1. 关联聚类标签到原数据

First, add your cluster assignments to your Data dataframe to easily group individuals:

# Add cluster labels to your original dataset
Data$cluster <- cl2.3

2. 方法一:计算个体到聚类质心的序列距离

We’ll use TraMineR’s seqcentroid to find the "central" sequence for each cluster, then calculate the Optimal Matching (OM) distance between each individual and their cluster’s centroid. This distance directly quantifies how far an individual is from the cluster’s core:

library(TraMineR)

# Calculate centroid sequence for each cluster
cluster_centroids <- lapply(sort(unique(cl2.3)), function(cluster_num) {
  # Filter sequences in the current cluster and compute its centroid
  seqcentroid(Data.seq_7[cl2.3 == cluster_num, ], 
              dist.matrix = dist.om2[cl2.3 == cluster_num, cl2.3 == cluster_num])
})
names(cluster_centroids) <- paste0("Cluster_", sort(unique(cl2.3)))

# Compute distance from each individual to their cluster's centroid
Data$centroid_distance <- sapply(1:nrow(Data), function(i) {
  current_cluster <- as.character(cl2.3[i])
  seqdist(Data.seq_7[i, ], cluster_centroids[[current_cluster]], 
          method = "OM", indel = 1, sm = submat2)
})

3. 方法二:计算个体到同聚类其他成员的平均距离

Another useful score is the average OM distance between an individual and all other members of its cluster. Higher values mean the individual is more dissimilar to the rest of the cluster:

# Calculate average distance to other cluster members (excluding self)
Data$avg_cluster_dissimilarity <- sapply(1:nrow(Data), function(i) {
  cluster_members <- which(cl2.3 == cl2.3[i])
  other_members <- cluster_members[cluster_members != i]
  mean(dist.om2[i, other_members])
})

如何分析极端值

Once you have these scores, identifying extreme individuals is straightforward:

  • Sort individuals within each cluster by centroid_distance or avg_cluster_dissimilarity (in descending order)
  • The top entries are your most extreme cases:
# Example: Get top 2 extreme individuals per cluster
library(dplyr)
extreme_individuals <- Data %>%
  group_by(cluster) %>%
  arrange(desc(avg_cluster_dissimilarity)) %>%
  slice_head(n = 2)

print(extreme_individuals)

补充说明

Note that the height values in your clusterward2 object represent the distance at which clusters were merged during the hierarchical process—these aren’t individual-level scores. The distance-based metrics we calculated above are far more relevant for identifying extreme individuals in your sequence clustering context.

内容的提问来源于stack exchange,提问作者Silvia Federzoni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 19:08:01