如何在R中提取无列名CSV的Study板块并整理为整洁格式?
在R中提取无列名CSV的Study板块并整理为整洁格式
核心思路
先定位Study板块的位置,提取该板块内容后,将拆分的表头和数据行分别合并,最终生成每行包含完整数据的整洁格式。
步骤1:读取并定位Study板块
首先读取无列名CSV,然后找到Study板块的起始与结束位置,提取该板块并移除空行:
# 读取无列名CSV,将空值识别为NA data1 <- read.csv("data.csv", header = FALSE, stringsAsFactors = FALSE, na.strings = "") # 定位Study板块的起始行 study_start <- which(data1$V1 == "Study") # 定位下一个板块的起始位置(若没有则取文件最后一行+1) next_section <- c(which(data1$V1 %in% c("Activities", "Sleep")), nrow(data1)+1) next_section_start <- min(next_section[next_section > study_start]) # 提取Study板块的所有行,并移除全空行 study_section <- data1[study_start:(next_section_start - 1), ] study_section <- study_section[!apply(study_section, 1, function(x) all(is.na(x))), ]
步骤2:合并表头与数据行
Study板块的表头分为两行,数据行也按两行一组拆分,需要分别合并:
# 合并两行表头为完整列名 header_part1 <- na.omit(study_section[2, ]) %>% as.character() header_part2 <- na.omit(study_section[3, ]) %>% as.character() full_header <- c(header_part1, header_part2) # 提取数据行(从第4行开始) raw_data_rows <- study_section[4:nrow(study_section), ] # 将奇数行与偶数行合并,生成完整数据行 clean_study_data <- cbind( raw_data_rows[seq(1, nrow(raw_data_rows), 2), ] %>% na.omit(), raw_data_rows[seq(2, nrow(raw_data_rows), 2), ] %>% na.omit() ) # 设置列名 colnames(clean_study_data) <- full_header
查看最终结果
执行以下代码即可得到符合要求的整洁格式:
# 打印Study标题及整理后的数据 cat("Study\n") print(clean_study_data, row.names = FALSE)
输出示例:
Study Date Calories_Burned Value1 Value2 Value3 Value4 1/6/2021 a b c d e 2/6/2021 a b c d e 3/6/2021 a b c d e
内容的提问来源于stack exchange,提问作者01200
相关产品推荐
相关产品推荐

