在R中读取包含多表头分隔数据框的文本文件问题求助
在R中读取包含多表头分隔数据框的文本文件问题求助
我有一个大文本文件,里面包含多个由表头行分隔的数据框,我正尝试将其读入R中。第一个表头行包含时间变量,我想根据时间变量分离这些数据框。数据示例如下:
data = c("** TIME: 41670", "** PROPERTY: Pressure", "** UNITS: psi", "< X > < Y > <Layer 1>", " 2106604.41 3434119.83 7952.25", " 2111884.40 3434119.83 7970.05", " 2037964.57 3439399.82 7658.27", " 2043244.56 3439399.82 7754", " 2048524.55 3439399.82 7828.24", " 2053804.53 3439399.82 7879.78", " 2059084.52 3439399.82 7914.57", " 2064364.50 3439399.82 7944.66", " 2069644.49 3439399.82 7974.44", " 2074924.48 3439399.82 7999.03", " 2080204.46 3439399.82 8014.14", " 2085484.46 3439399.82 8016.27", " 2090764.46 3439399.82 8005.63", "", "", "** TIME: 41670", "** PROPERTY: Pressure", "** UNITS: psi", "< X > < Y > <Layer 2>", " 2106604.41 3434119.83 8038.52", " 2111884.40 3434119.83 8066.89", " 2037964.57 3439399.82 7723.84", " 2043244.56 3439399.82 7821.79", " 2048524.55 3439399.82 7899.46", " 2053804.53 3439399.82 7955.23", " 2059084.52 3439399.82 7993.75", " 2064364.50 3439399.82 8026.08", " 2069644.49 3439399.82 8056.41", " 2074924.48 3439399.82 8080.33", " 2080204.46 3439399.82 8094.15", " 2085484.46 3439399.82 8095.07", " 2090764.46 3439399.82 8084.03", " 2096044.44 3439399.82 8068.33", " 2101324.41 3439399.82 8060.14", " 2106604.41 3439399.82 8073.08", " 2111884.40 3439399.82 8107.82", " 2117164.38 3439399.82 8145.84", " 2122444.37 3439399.82 8160.57" )
我目前用readLines来读取这个文本文件,理想的输出是一个列表,每个元素包含对应的时间戳和数据框,示例如下:
[[1]]$date [1] "2014-01-31" [[1]]$data X Y Layer1 1 2106604.41 3434119.83 7952.25 2 2111884.40 3434119.83 7970.05 3 2037964.57 3439399.82 7658.27 4 2043244.56 3439399.82 7754 [[2]]$date [1] "2014-01-31" [[2]]$data X Y Layer2 1 2106604.41 3434119.83 8038.52 2 2111884.40 3434119.83 8066.89 3 2037964.57 3439399.82 7723.84 4 2043244.56 3439399.82 7821.79
这是我尝试的代码:
data <- readLines("tmp.txt") # Initialize an empty list to store data frames dfs <- list() # Initialize variables current_time <- NULL current_df <- NULL property <- NULL # Loop through each line of the file for (line in data) { if (startsWith(line, "** TIME:")) { # Extract the time from the header line and convert to datetime current_time <- as.Date(as.numeric(trimws(sub("\\*{2}\\s+TIME:\\s+", "", line))), origin = "1899-12-30", format = "%Y-%m-%d") # Create a new data frame for the current time current_df <- data.frame() } else if (startsWith(line, "** PROPERTY:")) { next } else if (startsWith(line, "** UNITS:")) { next } else if (startsWith(line, "<")) { # Extract column names from header line 4 clean_header <- gsub("<|>", "", line) clean_header <- trimws(clean_header) col_names <- strsplit(clean_header, " ") col_names <- unlist(col_names) col_names <- col_names[col_names != ""] col_names[3] <- paste0(col_names[3], col_names[4]) col_names <- col_names[-4] } else if (!startsWith(line, "**")) { # Split the line by whitespace and create a new row in the data frame parts <- strsplit(line, "\\s+")[[1]] parts <- parts[parts != ""] current_df <- rbind(current_df, as.numeric(parts)) } else { # End of current data frame, store it in the list colnames(current_df) <- col_names dfs[[length(dfs) + 1]] <- list(date = current_time, data = current_df) current_df <- NULL } }
现在遇到的问题是:代码能生成current_df存储当前循环的最新数据框,但列名没有被添加进去;另外current_df也没有被保存到dfs列表中,会被新的current_df覆盖。
备注:内容来源于stack exchange,提问作者Bizzy
相关产品推荐
相关产品推荐

