基于DTW计算距离矩阵并用于聚类的技术咨询
Hey there! Let's break down how to tackle your DTW + clustering workflow step by step—no need to feel overwhelmed. I'll walk you through the R implementation since your data snippet looks like it's from R, which has fantastic tools for this kind of analysis.
Looking at your head(control[,1:4]) output, your rows are repeated samples per day (e.g., two Control_D1 entries), and columns are genes. DTW works best with single time series per feature (here, per gene), so first we need to aggregate repeated samples into one value per day per gene.
# Load required packages library(tidyverse) library(dtwclust) library(zoo) # Clean control data: aggregate repeated samples by day, sort chronologically control_clean <- control %>% rownames_to_column("sample_id") %>% # Extract day number from sample names mutate(day = str_extract(sample_id, "D\\d+") %>% str_remove("D") %>% as.integer()) %>% # Calculate mean for each gene per day (handle NAs if needed) group_by(day) %>% summarise(across(everything(), mean, na.rm = TRUE)) %>% # Ensure days are in order (D1 to D26) arrange(day) %>% column_to_rownames("day") %>% # Transpose: each row becomes a gene, each column is a day (D1-D26) t() %>% as.data.frame() # Repeat the same cleaning for treatment data treatment_clean <- treatment %>% rownames_to_column("sample_id") %>% mutate(day = str_extract(sample_id, "D\\d+") %>% str_remove("D") %>% as.integer()) %>% group_by(day) %>% summarise(across(everything(), mean, na.rm = TRUE)) %>% arrange(day) %>% column_to_rownames("day") %>% t() %>% as.data.frame() # Combine both datasets and add group labels for later visualization all_sequences <- rbind(control_clean, treatment_clean) %>% mutate(group = c(rep("control", nrow(control_clean)), rep("treatment", nrow(treatment_clean))))
Now we'll compute the pairwise DTW distances between all gene time sequences. The dtwclust package is optimized for this task, especially when clustering is the end goal.
# Extract only the time series columns (exclude the group label) ts_data <- all_sequences %>% select(-group) # Compute DTW distance matrix (adjust window.size based on your data's time continuity) dtw_dist_matrix <- dtw_dist(ts_data, window.size = 5)
Key Notes on DTW Parameters:
window.size: Limits how much the time series can be "warped" (e.g., a window of 5 allows aligning day 1 with days 1-5). Adjust based on how much temporal shift you expect in your data.- If you have missing values, use
zoo::na.approx(ts_data)to interpolate gaps before calculating distances—DTW doesn't handle missing data natively.
With your distance matrix ready, you can use standard clustering methods that work with distance inputs. Below are two common approaches:
Option 1: Hierarchical Clustering (Great for Visualization)
library(dendextend) # Run hierarchical clustering hc <- hclust(dtw_dist_matrix, method = "ward.D2") # Visualize the dendrogram with colored labels for control/treatment groups hc_dend <- as.dendrogram(hc) hc_dend <- color_labels(hc_dend, labels = all_sequences$group, col = c("steelblue", "tomato")) plot(hc_dend, cex = 0.7, main = "Hierarchical Clustering of Gene Time Series (DTW Distance)")
Option 2: K-Means Clustering (Great for Assigning Cluster Labels)
# Convert distance matrix to a format compatible with k-means dtw_dist_mat <- as.matrix(dtw_dist_matrix) # Run k-means (choose k based on your data, e.g., 4 clusters) km_clusters <- kmeans(dtw_dist_mat, centers = 4) # View cluster distribution across control/treatment groups table(Cluster = km_clusters$cluster, Group = all_sequences$group)
内容的提问来源于stack exchange,提问作者Zizogolu

