加载TSV文件时维度不匹配报错的排查与解决求助
我需要将一个TSV文件导入R包使用,执行datafile <- "omics-data.txt"指定路径正常,但调用包函数时报错:
Error in dimnames(x) <- dn :
length of 'dimnames' [1] not equal to array extent
排查后确认问题出在文件本身,用read.table("omics-data.txt", sep = "\t", header = FALSE)加载时也报错:
Error in scan(file = file, what = what, sep = sep, quote = quote, dec = dec, :
line 339 did not have 28 elements
检查第339行看起来包含28个元素,但删除该行后错误又指向其他行。
数据结构
glimpse(data)输出:
Rows: 3,145 Columns: 28 $ id <chr> "Cholestanyl glucoside_1", "DG(16:0/16:0/0:0)_1", "Uracil_1", "1-(8-[3]-ladderane-octanyl)-2-(8-[3]-ladderane-octanyl)-sn-glycerol_1", "CerP(d18:1/8:0)_1", "L-Serine_1", "Myristic a… $ `74.8591` <dbl> -1.545185800, 0.740397700, 0.426691120, 0.236691010, 0.036180700, -1.873424450, -3.519595738, -1.158267900, -0.818563280, 0.407870000, -0.124774800, 0.745390640, 1.265679227, 0.4942… $ `78.5399` <dbl> -1.618735800, -1.555572500, 0.449842740, 0.009369260, -2.877068370, -1.651203510, -0.008264896, 0.762969100, 0.024828640, 1.929355300, 0.181821900, 1.354894500, -2.368993573, 0.2272…
head(data)输出:
# A tibble: 6 × 28 id `74.8591` `78.5399` `58.2187` `60.5529` `27.6748` `61.2103` `79.7051` `80.212` `30.0669` `60.9525` `81.0258` `99.1194` `43.9243` `66.9193` `71.1771` `97.8946` `31.9547` `33.7501` `80.9173` <chr> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> 1 Cholestan… -1.55 -1.62 -1.57 0.660 0.701 -1.48 0.697 0.469 -1.45 0.597 0.475 0.576 0.594 0.642 -1.54 0.815 0.494 0.661 0.672 2 DG(16:0/1… 0.740 -1.56 0.642 -1.30 0.625 0.607 -1.35 0.520 0.734 0.686 0.590 -0.706 -1.48 0.731 -1.55 0.884 -1.39 0.727 -1.42
数据说明:第一列为ID,第2至28列为27个样本,列名为关联数字,所有数据已做对数转换,无NA值。尝试多种加载方式仍未解决,求帮助。
定位异常行的隐藏问题:用
readLines读取所有行,统计每行的制表符数量(28列需要27个分隔符),找出数量不符的行:lines <- readLines("omics-data.txt") tab_counts <- sapply(lines, function(x) length(gregexpr("\t", x)[[1]])) which(tab_counts != 27)使用更宽容的加载函数:用
readr::read_tsv替代基础函数,它会给出详细的错误提示,还能通过problems()查看所有格式异常的行:library(readr) df <- read_tsv("omics-data.txt") problems(df)强制指定列类型:确保ID列被识别为字符型,避免解析错误:
df <- read.table("omics-data.txt", sep = "\t", header = TRUE, colClasses = c("character", rep("numeric", 27)))手动清理文件:用支持显示特殊字符的编辑器(如Notepad++)打开文件,开启「显示所有字符」,检查异常行是否有隐藏换行符(如
\r)或多余制表符,直接修正。
内容的提问来源于stack exchange,提问作者Sarah Green

