You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中提取聚类结果中含特定数量基因的簇内基因信息

基因聚类及目标簇基因提取实现方案

1. 数据准备

先将Dropbox上的Excel文件下载到本地,再用R读取并预处理:

# 加载依赖包
library(readxl)

# 读取数据,第一列为基因特征编号
gene_df <- read_excel("Sums.xlsx")
# 将基因编号设为行名,方便后续聚类后对应基因
rownames(gene_df) <- gene_df[[1]]
# 仅保留症状标记列用于距离计算
gene_features <- gene_df[, -1]

2. 计算距离与层次聚类

按照要求使用binary方法计算距离,再完成层次聚类并指定高度切割:

# 计算二进制距离矩阵
dist_matrix <- dist(gene_features, method = "binary")

# 执行层次聚类(默认采用ward.D2方法,可根据数据特性更换为complete/average等)
hclust_result <- hclust(dist_matrix)

# 按高度h=0.35切割聚类树,得到每个基因的簇分配结果
cluster_assignments <- cutree(hclust_result, h = 0.35)

3. 筛选并提取含7个基因的簇

统计各簇的基因数量,找出符合条件的簇并输出对应基因编号:

# 统计每个簇的基因数量
cluster_counts <- table(cluster_assignments)

# 筛选出恰好包含7个基因的簇编号
target_cluster_ids <- names(cluster_counts[cluster_counts == 7])

# 逐个输出目标簇的基因列表
for (cid in target_cluster_ids) {
  cat("簇", cid, "的基因编号:\n")
  cat(rownames(gene_df)[cluster_assignments == cid], sep = "\n")
  cat("\n")
}

注意事项

  • 若无需下载文件,可使用rdrop2包直接从Dropbox读取,但需先完成Dropbox的API授权配置
  • 若聚类结果与预期的21个7基因簇有出入,可调整聚类方法或微调切割高度h的数值

内容的提问来源于stack exchange,提问作者Marina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 18:58:33