You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pivot_wider处理DepMap蛋白质组数据遇错误求解决方案

DepMap蛋白质组数据宽表转换错误解决方案

问题描述

从DepMap下载蛋白质组数据后,尝试将depmap_group_id设为行名、protein_id设为列名,使用pivot_wider转换时遇到两个错误:

  • 初始转换报错:Error in vec_init(): ! nmust be a single number, not an integerNA``
  • 缩小数据集至原大小20%后,仍报错:Error: cannot allocate vector of size 10.4 Gb

解决方法

1. 清理数据中的NA值

第一个错误通常是关键列存在NA导致vec_init无法确定向量长度,先检查并清理:

# 检查各关键列的NA数量
table(is.na(proteomic_dep_g_ID$depmap_group_id))
table(is.na(proteomic_dep_g_ID$protein_id))
table(is.na(proteomic_dep_g_ID$protein_expression))

# 移除含NA的行(若需保留可改用fill填充,根据业务需求调整)
proteomic_dep_g_ID_clean <- proteomic_dep_g_ID %>%
  drop_na(depmap_group_id, protein_id, protein_expression)

2. 使用内存效率更高的data.table::dcast替代pivot_wider

pivot_wider处理超大数据集时内存占用极高,data.table的dcast在内存管理上更高效,适合这类场景:

library(data.table)

# 转换为data.table格式
setDT(proteomic_dep_g_ID_clean)

# 执行宽表转换,fun.aggregate=mean对应原需求的values_fn=mean
proteomic_dep_g_ID2 <- dcast(proteomic_dep_g_ID_clean, 
                             depmap_group_id ~ protein_id, 
                             value.var = "protein_expression",
                             fun.aggregate = mean)

3. 分批处理缓解内存压力

如果内存依然不足,可按depmap_group_id拆分批次转换后合并:

# 获取唯一的depmap_group_id并拆分批次
group_ids <- unique(proteomic_dep_g_ID_clean$depmap_group_id)
batch_size <- 100  # 根据自身内存情况调整批次大小
batches <- split(group_ids, ceiling(seq_along(group_ids)/batch_size))

# 分批转换并合并结果
result_list <- lapply(batches, function(batch) {
  sub_data <- proteomic_dep_g_ID_clean[depmap_group_id %in% batch]
  dcast(sub_data, depmap_group_id ~ protein_id, 
        value.var = "protein_expression", fun.aggregate = mean)
})

proteomic_dep_g_ID2 <- rbindlist(result_list, fill = TRUE)

4. 检查并处理重复数据

重复的depmap_group_id+protein_id组合会增加内存负担,提前检查并处理:

# 找出重复的组合
duplicate_records <- proteomic_dep_g_ID_clean %>%
  count(depmap_group_id, protein_id) %>%
  filter(n > 1)

# 若重复为冗余数据,可去重;若为合理重复则保留
proteomic_dep_g_ID_clean <- proteomic_dep_g_ID_clean %>%
  distinct(depmap_group_id, protein_id, .keep_all = TRUE)

内容的提问来源于stack exchange,提问作者mei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 04:53:33