You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言hclust函数疑似Bug:数据框子集化不当致标签丢失

Issue with hclust() Losing Labels After Improper Data Frame Subsetting

Hey everyone, I recently noticed a frustrating quirk when using R's hclust() function for clustering: if you subset your data frame the "wrong" way, the resulting clustering object loses its row labels entirely. Let me walk through the problem, why it happens, and how to fix it.

Step 1: Test Data Setup

First, let's create our sample data frame to reproduce the issue:

set.seed(1234)
x <- rnorm(12, mean = rep(1:3, each = 4), sd = 0.2)
y <- rnorm(12, mean = rep(c(1,2,1), each = 4), sd = 0.2)
z <- as.factor(sample(c("A","B"), 12, replace=T))
df <- data.frame(x=x, y=y, z=z)

# Plot to visualize the data
plot(df$x, df$y, col=z, pch=19, cex=2)

Step 2: The Problem - Improper Subsetting

Let's say we want to cluster only the rows where z == "A". If we subset using this approach, we'll lose labels in the hclust result:

# Bad subsetting: drops meaningful labels
df_sub_bad <- df[df$z == "A", c("x", "y")]

# Compute distance matrix and cluster
dist_bad <- dist(df_sub_bad)
hc_bad <- hclust(dist_bad)

# Check the labels - they're gone!
hc_bad$labels
# Output: NULL

Why This Happens

When you subset a data frame with df[df$z == "A", c("x", "y")], R resets the row names to sequential integers (1, 2, ...). The dist() function doesn't recognize these default integer row names as meaningful labels, so it doesn't pass them along to the hclust() object. That's why we end up with NULL for labels.

Step 3: Fixes - Proper Subsetting Methods

Here are two reliable ways to keep the labels intact:

Method 1: Use drop=FALSE During Subsetting

Adding drop=FALSE ensures the subset remains a data frame (critical if you end up with a single row!) and preserves the original row names:

# Good subsetting: retain original row names
df_sub_good1 <- df[df$z == "A", c("x", "y"), drop = FALSE]

# Compute distance and cluster
dist_good1 <- dist(df_sub_good1)
hc_good1 <- hclust(dist_good1)

# Labels are preserved!
hc_good1$labels
# Output: original row indices from the parent data frame

Method 2: Convert Subset to a Matrix First

Converting the subset to a matrix before calculating distance forces R to retain row names as labels:

# Convert subset to matrix
df_sub_matrix <- as.matrix(df[df$z == "A", c("x", "y")])

dist_good2 <- dist(df_sub_matrix)
hc_good2 <- hclust(dist_good2)

# Labels are here too!
hc_good2$labels

Bonus: Verify with a Dendrogram

Plotting the dendrogram will confirm labels are present and correctly mapped:

plot(hc_good1, main = "Clustering with Preserved Labels")

Key Takeaway

The main issue is that default data frame subsetting can overwrite meaningful row names with generic integers, which dist() ignores. Using drop=FALSE or converting to a matrix ensures your row labels stick around through the clustering process.

内容的提问来源于stack exchange,提问作者alebj88

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:26:44