You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TraMineR工具:预测新离散状态序列的类别归属

Predicting Cluster Membership for New Sequences with TraMineR

Got it, let's walk through how to assign new discrete state sequences to existing clusters using TraMineR, building on your precomputed dissimilarity matrix D and clustering results. Here's a practical, step-by-step breakdown with R code examples:

1. Align New Sequences with Training Data Structure

First, you need to make sure your new sequences match the training data's alphabet, state labels, and encoding exactly—otherwise your distance calculations will be meaningless. Use seqdef to standardize them:

# Assume your preprocessed training sequence object is named train.seq
new_seq_raw <- # Your new sequence data (e.g., a data frame/vector of state strings)
new.seq <- seqdef(
  data = new_seq_raw,
  alphabet = alphabet(train.seq),  # Reuse the training set's state alphabet
  labels = labels(train.seq),      # Match state labels from training
  states = states(train.seq),      # Keep state encoding consistent
  missing = NA                     # Handle missing values the same way as training
)

2. Calculate Distances from New Sequences to Training Data

Compute the dissimilarity between each new sequence and every sequence in your training set. Critical: Use the exact same dissimilarity method and parameters you used to create matrix D (e.g., Optimal Matching with your cost matrix, LCS, Hamming distance):

# Example: If you used Optimal Matching (OM) with a substitution cost matrix 'cost.mat'
new_distances <- seqdist(
  seq1 = new.seq,
  seq2 = train.seq,
  method = "OM",  # Match your original method (e.g., "LCS", "HAM")
  sm = cost.mat   # Reuse your original substitution cost matrix (if applicable)
)

# new_distances will be a 1 x N matrix (N = number of training sequences)

3. Predict Cluster Membership

The method depends on the type of clustering you ran on your training data:

Case 1: K-Means Clustering

If you used seqclust(method = "kmeans"), your result object includes cluster centroids. Calculate the distance from the new sequence to each centroid, then assign it to the closest cluster:

# Assume your k-means clustering result is stored in km_clust
km_clust <- seqclust(train.seq, method = "kmeans", diss = D, k = 3)

# Calculate distance from new sequence to each centroid
centroid_distances <- seqdist(
  seq1 = new.seq,
  seq2 = km_clust$centers,
  method = "OM",
  sm = cost.mat
)

# Assign to the cluster with the smallest distance
predicted_cluster <- which.min(centroid_distances)

Case 2: Hierarchical Clustering

For hierarchical clustering (e.g., agnes or diana paired with seqdist), a simple k-Nearest Neighbors (k-NN) vote works best:

  1. Find the k closest training sequences to your new sequence
  2. Assign the most common cluster label among those neighbors
# Assume your training cluster labels are stored in train_clusters (from cutree())
train_clusters <- cutree(hierarchical_clust_obj, k = 3)

# Choose k (e.g., 5 nearest neighbors—adjust based on your data size)
k <- 5

# Get indices of the k closest training sequences
nearest_indices <- order(new_distances)[1:k]

# Pull their cluster labels
nearest_clusters <- train_clusters[nearest_indices]

# Vote for the most frequent cluster
predicted_cluster <- names(sort(table(nearest_clusters), decreasing = TRUE))[1]

Quick Pro Tips

  • Always verify state encoding consistency (e.g., if "S" = Single in training, don’t use "Single" in new data).
  • For large training sets, k-NN with a small k (3-7) is efficient and avoids recomputing the full dissimilarity matrix.
  • If you want to avoid k-NN for hierarchical clustering, you could append the new sequence to training data and re-run clustering—but this is computationally heavy for big datasets.

内容的提问来源于stack exchange,提问作者Maxim.K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:55:36