R语言:读取不规则文件中USAGE部分为DataFrame的问题
Got it, let's tackle this problem step by step. I've run into similar issues with structured text files in R before, so here's what I'd recommend:
核心问题分析
Your variable-based skip/nrows failure is almost certainly due to off-by-one errors in how you're calculating those values (e.g., including the separator line itself in your data range, or miscalculating the number of rows between separators). Instead of relying on read.table's skip/nrows directly, a more robust approach is to first load all lines into memory, locate your separators, then extract each component precisely.
推荐解决方案:先读全行再分割
This method gives you full visibility into where each section starts and ends, making debugging way easier. Here's how to do it:
1. 读取所有行到向量
First, load the entire file into a character vector so you can inspect and index lines easily:
all_lines <- readLines("your_file_path.txt")
2. 定位各个分隔符的行号
Use grep to find the exact positions of your section markers. Add ^ and $ to ensure you match exact lines (adjust if your separators have extra whitespace with ^\\s* and \\s*$):
# 匹配精确的分隔符行(如果分隔符前后有空格,改成 "^\\s*SOURCE\\s*$" 这类) source_pos <- grep("^SOURCE$", all_lines)[1] # 取第一个匹配(确保每个分隔符只出现一次) story_pos <- grep("^STORY$", all_lines)[1] usage_pos <- grep("^USAGE$", all_lines)[1] dataset_pos <- grep("^DATASET$", all_lines)[1]
3. 提取每个组件为独立R对象
Now you can slice the all_lines vector to get each section, then convert to the appropriate object type:
SOURCE 部分(文本类)
source_content <- paste(all_lines[(source_pos + 1):(story_pos - 1)], collapse = "\n") # 如果是结构化数据,可根据情况转成数据框,否则保留为字符向量/字符串
STORY 部分(文本类)
story_content <- paste(all_lines[(story_pos + 1):(usage_pos - 1)], collapse = "\n")
USAGE 部分(固定4列表格)
Extract the relevant lines first, then use read.table with the text parameter to parse directly from the character vector:
usage_data_lines <- all_lines[(usage_pos + 1):(dataset_pos - 1)] usage_df <- read.table( text = usage_data_lines, col.names = c("Col1", "Col2", "Col3", "Col4"), # 替换成你的实际列名 header = FALSE, # 如果USAGE部分有表头,改成TRUE并调整col.names stringsAsFactors = FALSE # 按需设置 )
DATASET 部分(表格类)
dataset_data_lines <- all_lines[(dataset_pos + 1):length(all_lines)] dataset_df <- read.table( text = dataset_data_lines, # 根据实际列数和格式添加参数:col.names, sep, header等 stringsAsFactors = FALSE )
如果坚持用read.table的skip/nrows参数
If you want to stick with your original approach, double-check your variable calculations. For the USAGE section:
skipshould be the line number of theUSAGEseparator itself (sinceskiptellsread.tableto ignore the first N lines, so data starts at N+1)nrowsshould bedataset_pos - usage_pos - 1(subtract 1 to exclude both theUSAGEandDATASETseparator lines)
Example code:
usage_df <- read.table( "your_file_path.txt", skip = usage_pos, nrows = (dataset_pos - usage_pos - 1), col.names = c("Col1", "Col2", "Col3", "Col4"), header = FALSE )
Just make sure usage_pos and dataset_pos are the correct line numbers (use print(usage_pos) to verify—common mistakes include counting from 0 instead of 1, or matching multiple separator lines).
额外小贴士
- Always verify your separator positions with
print(c(source_pos, story_pos, usage_pos, dataset_pos))to ensure they're in the right order and unique. - If your file has inconsistent line endings or encoding, add
encoding = "UTF-8"(or appropriate encoding) toreadLines.
内容的提问来源于stack exchange,提问作者ProfessorE

