如何提取矩阵中行与列名称前7字符相同的数值
解决方案
核心逻辑
每个矩形对应4个前缀(前7字符)相同的点,4个点的不重复两两组合恰好是6组,因此只需筛选行名与列名前7字符一致且行名小于列名(避免重复提取A-B和B-A)的距离值,再按前缀分组即可。
R代码实现
假设你的距离矩阵名为dist_mat,执行以下步骤:
- 提取所有行/列名的前7个字符:
row_prefix <- substr(rownames(dist_mat), 1, 7) col_prefix <- substr(colnames(dist_mat), 1, 7)
- 创建筛选掩码,锁定符合条件的位置:
filter_mask <- row_prefix == col_prefix & rownames(dist_mat) < colnames(dist_mat)
- 提取目标值并按矩形前缀分组:
rect_distances <- split(dist_mat[filter_mask], row_prefix[filter_mask])
样本验证示例
用模拟数据测试代码效果:
# 生成模拟距离矩阵 sample_ids <- c("RECT001_P1", "RECT001_P2", "RECT001_P3", "RECT001_P4", "RECT002_P1", "RECT002_P2", "RECT002_P3", "RECT002_P4") dist_mat <- matrix(runif(64), nrow=8, dimnames=list(sample_ids, sample_ids)) diag(dist_mat) <- 0 # 清除自身距离的0值 # 执行提取逻辑 row_prefix <- substr(rownames(dist_mat), 1, 7) col_prefix <- substr(colnames(dist_mat), 1, 7) filter_mask <- row_prefix == col_prefix & rownames(dist_mat) < colnames(dist_mat) rect_distances <- split(dist_mat[filter_mask], row_prefix[filter_mask]) # 检查每个矩形的距离数量 lapply(rect_distances, length) # 输出应为每个分组对应6个值
内容的提问来源于stack exchange,提问作者KimChiya
相关产品推荐
相关产品推荐

