R语言处理质谱数据:按阈值过滤后分组计算肽段平均强度
解决方案
方案1:统一使用5000作为强度截断值
用dplyr包的分组聚合逻辑实现,代码如下:
# 加载依赖包 library(dplyr) # 你的测试数据 test_Data <- structure(list(UNIPROT = structure(c(2L, 2L, 2L, 2L, 2L, 2L, 2L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L), .Label = c("A8DUK4", "P08032", "P15508"), class = "factor"), Intensity = c(16926.19, 36738.94, 2203.22, 5338.85, 133.44, 27991.35, 29505.84, 201.4695, 47469.09, 24841.01, 4546.9, 22805.69, 18494.71, 28805.99, 68220.65, 90526.29, 63259.19, 44492.48, 65497.13, 40704.81, 334874.1, 38702.87, 300135)), class = "data.frame", row.names = c(NA, -23L)) # 过滤+求平均 result <- test_Data %>% group_by(UNIPROT) %>% filter(Intensity > 5000) %>% summarise(mean_Intensity = mean(Intensity))
运行后得到的结果如下:
| UNIPROT | mean_Intensity |
|---|---|
| A8DUK4 | 116279.2 |
| P08032 | 23220.23 |
| P15508 | 28483.3 |
方案2:为不同UNIPROT设置自定义阈值
可以先建立一个阈值映射表,再关联原数据做过滤计算:
# 建立自定义阈值表,可根据需求修改对应数值 threshold_df <- tibble( UNIPROT = c("P08032", "P15508", "A8DUK4"), cut_off = c(5000, 4000, 30000) # 分别对应三个蛋白的阈值 ) # 关联后过滤计算 result_custom <- test_Data %>% left_join(threshold_df, by = "UNIPROT") %>% filter(Intensity > cut_off) %>% group_by(UNIPROT) %>% summarise(mean_Intensity = mean(Intensity))
如果某类UNIPROT过滤后没有符合条件的行,结果中会返回NA,你可以根据需要在mean()函数里加na.rm = TRUE参数,或者后续过滤掉NA行即可。
内容的提问来源于stack exchange,提问作者PesKchan
相关产品推荐
相关产品推荐

