如何拆分长度可变的字符串并转换为带NA填充的DataFrame
解决方案
针对你的需求,这里提供两种R语言实现方法,都能将字符串向量拆分成目标格式的DataFrame,缺失值自动填充NA:
方法一:Base R 原生实现
# 定义原始数据 mydat <- c("1-2 1-1", "1-2 1-1 3-3", "1-1 2-1 4-1", "1-1") # 1. 按空格拆分每个字符串为x-y片段列表 split_fragments <- strsplit(mydat, " ") # 2. 将每个x-y片段拆分为数字,扁平化处理 flat_nums <- lapply(split_fragments, function(frag) as.numeric(unlist(strsplit(frag, "-")))) # 3. 计算最长的数字序列长度,给短序列补NA max_length <- max(sapply(flat_nums, length)) filled_nums <- lapply(flat_nums, function(x) c(x, rep(NA, max_length - length(x)))) # 4. 转换为DataFrame并命名列 newdat <- as.data.frame(do.call(rbind, filled_nums)) colnames(newdat) <- c("one", "two", "three", "four", "five", "six") # 查看结果 newdat
方法二:Tidyverse 工具链实现
如果习惯用tidyverse的语法,代码更简洁直观:
library(tidyverse) # 定义原始数据 mydat <- c("1-2 1-1", "1-2 1-1 3-3", "1-1 2-1 4-1", "1-1") newdat <- tibble(raw_str = mydat) %>% # 按空格拆分每行,生成多行数据 separate_rows(raw_str, sep = " ") %>% # 将x-y拆分为两个数值列 separate(raw_str, into = c("num1", "num2"), sep = "-", convert = TRUE) %>% # 为原始每行数据添加分组ID group_by(orig_id = rep(1:length(mydat), each = n()/length(mydat))) %>% # 生成目标列名的映射 mutate(col_name = rep(c("one", "two"), n())) %>% ungroup() %>% # 转换为宽格式 pivot_wider(names_from = col_name, values_from = c(num1, num2)) %>% # 清理列名并按顺序排列 rename_with(~str_remove(., "num\\d+_")) %>% select(one, two, three, four, five, six) # 查看结果 newdat
两种方法运行后都会得到你需要的newdat格式,缺失位置自动填充NA。
内容的提问来源于stack exchange,提问作者Joe
相关产品推荐
相关产品推荐

