使用TCGAbiolinks下载TCGA-DLBC数据时R报向量过长错误怎么办?
TCGAbiolinks下载TCGA-DLBC数据报错的解决办法
问题描述
使用TCGAbiolinks查询TCGA-DLBC的RNA-seq比对BAM文件时,getResults(query_TCGA)可正常返回结果,但执行GDCdownload(query_TCGA)时出现以下报错:
Downloading data for project TCGA-DLBC
GDCdownload will download: 11.806323702 GB
The total size of files is big. We will download files in chunks
At least one of the chunks download was not correct. We will retry
Error in 0:ceiling(nrow(manifest)/step - 1) :
result would be too long a vector
解决办法
1. 手动调整分块下载步长
报错源于默认分块逻辑生成的序列向量超出R的长度限制,可通过指定chunks.per.download参数缩小分块大小:
GDCdownload(query_TCGA, method = "api", chunks.per.download = 1)
可根据实际文件数量调整chunks.per.download的值(如1或2),避免生成过长的分块索引向量。
2. 使用GDC官方客户端下载
若R包的下载逻辑持续异常,可导出manifest文件后用GDC CLI工具下载:
- 导出manifest文件:
write.table(getManifest(query_TCGA), "TCGA-DLBC_manifest.txt", sep = "\t", row.names = FALSE, quote = FALSE)
- 用GDC CLI执行下载(需提前安装GDC客户端):
gdc-client download -m TCGA-DLBC_manifest.txt
3. 更新TCGAbiolinks到最新版本
旧版本TCGAbiolinks可能存在分块下载的bug,更新到最新版可修复部分已知问题:
if (!require("BiocManager", quietly = TRUE)) install.packages("BiocManager") BiocManager::install("TCGAbiolinks")
4. 单个文件单独下载
检查查询结果中的文件数量,若存在多个文件,可逐个单独下载:
- 提取所有文件ID:
file_ids <- getResults(query_TCGA)$file_id
- 循环下载每个文件:
for (id in file_ids) { query_single <- GDCquery(project = 'TCGA-DLBC', file.id = id) GDCdownload(query_single) }
内容的提问来源于stack exchange,提问作者Mirue Kang
相关产品推荐
相关产品推荐

