You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将含交错列名的文本文件读取到R语言中?

在R中读取列名分布在多行的文本文件

问题描述

有如下格式的文本文件:

* +---------------- Station Code
* |    +----------- Schedule Arrival Day
* |    |  +-------- Schedule Arrival Time
* |    |  |     +----- Schedule Departure Day
* |    |  |     |  +-- Schedule Departure Time
* |    |  |     |  |     +------------- Actual Arrival Time
* |    |  |     |  |     |     +------- Actual Departure Time
* |    |  |     |  |     |     |     +- Comments
* V    V  V     V  V     V     V     V
* NOL  *  *     1  900A  *     900A  Departed:  On time.
* SCH  *  *     1  1030A *     1039A Departed:  9 minutes late.
* NIB  *  *     1  1156A *     1159A Departed:  3 minutes late.
* LFT  *  *     1  1224P *     1228P Departed:  4 minutes late.
* LCH  *  *     1  155P  *     155P  Departed:  On time.

文件的列名分布在不同行中,需要将其读取到R语言中。

解决步骤

1. 读取并预处理所有行

先用readLines读取文件全部内容,再清理每行开头的*和冗余空格:

# 替换为你的文件实际路径
file_path <- "train_schedule.txt"
lines <- readLines(file_path)
# 移除每行开头的*及后续空格
lines <- gsub("^\\*\\s*", "", lines)

2. 提取列名

从带+的行里提取每个列的名称:

# 筛选列名定义行(包含"+"的行)
col_def_lines <- lines[grepl("\\+", lines)]
# 提取每行末尾的列名并去除前后空格
col_names <- sapply(col_def_lines, function(x) trimws(sub(".*\\s", "", x)))

3. 确定列的分割位置

利用标记列位置的V行,生成各列的起止索引:

# 找到列位置标记行(包含多个V的行)
pos_line <- lines[grepl("V\\s+V", lines)]
# 获取每个V的起始位置
pos <- gregexpr("V", pos_line)[[1]]
# 补充行尾位置作为最后一列的结束点
pos <- c(pos, nchar(pos_line))
# 生成各列的起止范围数据框
col_ranges <- data.frame(start = pos[-length(pos)], end = pos[-1] - 1)

4. 提取并整理数据行

筛选出数据行,按列分割后转换为标准数据框:

# 筛选数据行(排除列定义和位置标记行)
data_lines <- lines[!grepl("\\+|V\\s+V", lines)]

# 按列分割每行数据并去除空格
data_list <- lapply(data_lines, function(line) {
  sapply(1:nrow(col_ranges), function(i) {
    trimws(substr(line, col_ranges$start[i], col_ranges$end[i]))
  })
})

# 转换为数据框并设置列名
train_df <- as.data.frame(do.call(rbind, data_list), stringsAsFactors = FALSE)
colnames(train_df) <- col_names

5. 优化数据类型(可选)

根据需求转换列的类型,比如时间、数值型:

# 先安装并加载lubridate包(如果未安装:install.packages("lubridate"))
library(lubridate)

# 将调度日转为数值型
train_df$`Schedule Departure Day` <- as.numeric(train_df$`Schedule Departure Day`)

# 转换时间列为标准时间格式
train_df$`Schedule Departure Time` <- parse_date_time(train_df$`Schedule Departure Time`, "%I%M%p")
train_df$`Actual Departure Time` <- parse_date_time(train_df$`Actual Departure Time`, "%I%M%p")

最终效果

运行完上述代码后,train_df就是整理好的结构化数据框,包含完整列名和清洗后的数据。

内容的提问来源于stack exchange,提问作者mjc0203

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 22:02:05