You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于DTW计算距离矩阵并用于聚类的技术咨询

Hey there! Let's break down how to tackle your DTW + clustering workflow step by step—no need to feel overwhelmed. I'll walk you through the R implementation since your data snippet looks like it's from R, which has fantastic tools for this kind of analysis.

Step 1: Data Preparation & Cleaning

Looking at your head(control[,1:4]) output, your rows are repeated samples per day (e.g., two Control_D1 entries), and columns are genes. DTW works best with single time series per feature (here, per gene), so first we need to aggregate repeated samples into one value per day per gene.

# Load required packages
library(tidyverse)
library(dtwclust)
library(zoo)

# Clean control data: aggregate repeated samples by day, sort chronologically
control_clean <- control %>%
  rownames_to_column("sample_id") %>%
  # Extract day number from sample names
  mutate(day = str_extract(sample_id, "D\\d+") %>% str_remove("D") %>% as.integer()) %>%
  # Calculate mean for each gene per day (handle NAs if needed)
  group_by(day) %>%
  summarise(across(everything(), mean, na.rm = TRUE)) %>%
  # Ensure days are in order (D1 to D26)
  arrange(day) %>%
  column_to_rownames("day") %>%
  # Transpose: each row becomes a gene, each column is a day (D1-D26)
  t() %>%
  as.data.frame()

# Repeat the same cleaning for treatment data
treatment_clean <- treatment %>%
  rownames_to_column("sample_id") %>%
  mutate(day = str_extract(sample_id, "D\\d+") %>% str_remove("D") %>% as.integer()) %>%
  group_by(day) %>%
  summarise(across(everything(), mean, na.rm = TRUE)) %>%
  arrange(day) %>%
  column_to_rownames("day") %>%
  t() %>%
  as.data.frame()

# Combine both datasets and add group labels for later visualization
all_sequences <- rbind(control_clean, treatment_clean) %>%
  mutate(group = c(rep("control", nrow(control_clean)), rep("treatment", nrow(treatment_clean))))
Step 2: Calculate DTW Distance Matrix

Now we'll compute the pairwise DTW distances between all gene time sequences. The dtwclust package is optimized for this task, especially when clustering is the end goal.

# Extract only the time series columns (exclude the group label)
ts_data <- all_sequences %>% select(-group)

# Compute DTW distance matrix (adjust window.size based on your data's time continuity)
dtw_dist_matrix <- dtw_dist(ts_data, window.size = 5)

Key Notes on DTW Parameters:

  • window.size: Limits how much the time series can be "warped" (e.g., a window of 5 allows aligning day 1 with days 1-5). Adjust based on how much temporal shift you expect in your data.
  • If you have missing values, use zoo::na.approx(ts_data) to interpolate gaps before calculating distances—DTW doesn't handle missing data natively.
Step 3: Cluster Analysis with DTW Distances

With your distance matrix ready, you can use standard clustering methods that work with distance inputs. Below are two common approaches:

Option 1: Hierarchical Clustering (Great for Visualization)

library(dendextend)

# Run hierarchical clustering
hc <- hclust(dtw_dist_matrix, method = "ward.D2")

# Visualize the dendrogram with colored labels for control/treatment groups
hc_dend <- as.dendrogram(hc)
hc_dend <- color_labels(hc_dend, labels = all_sequences$group, col = c("steelblue", "tomato"))
plot(hc_dend, cex = 0.7, main = "Hierarchical Clustering of Gene Time Series (DTW Distance)")

Option 2: K-Means Clustering (Great for Assigning Cluster Labels)

# Convert distance matrix to a format compatible with k-means
dtw_dist_mat <- as.matrix(dtw_dist_matrix)

# Run k-means (choose k based on your data, e.g., 4 clusters)
km_clusters <- kmeans(dtw_dist_mat, centers = 4)

# View cluster distribution across control/treatment groups
table(Cluster = km_clusters$cluster, Group = all_sequences$group)

内容的提问来源于stack exchange,提问作者Zizogolu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:55:27