如何在R中将一列拆分为多列,使同列存储相似值
拆分R数据集中的单列内容为多列(相似内容归为同一列,缺失填NA)
我是R语言新手,现在有一个数据集的单列内容,想要把它拆分成多列,要求拆分后的每列存储相似内容,某行没有对应内容时填充NA。以下是该列的10行数据:
extract_df$Res <- c("\n- Provided disclaimer, Disconnected", "\n- confirmed the number,\n- pin was verified,\n- Explained reactivation,\n- confirmed the zip-code,\n - process explained,\n- changed the billing cycle,\n- Bill Amount : 60 dollars", "\n- confirmed the number", "\n- confirmed the number,\n- pin was verified,\n- Confirmed that the line is suspended due to non-payment", "\n- confirmed the number,\n- pin was verified,\n- Explained reactivation,\n- Activation fee charged,\n- Confirmed that the line is suspended due to non-payment,\n- $4 processing fee waived,\n- informed the customer about the payment and changing billing cycle,\n- changed the billing cycle,\n- manually processed the payment,\n- Payment completed successfully,\n- Due Date : next month,\n- Promised back", "\n- confirmed the number,\n- pin was verified,\n- Requested for sim number", "\n- confirmed the number,\n- Inquired about signal or network status", "\n- confirmed the number,\n- Explained reactivation,\n- Activation fee charged,\n- Confirmed that the line is suspended due to non-payment,\n- confirmed the zip-code,\n- $4 processing fee waived,\n - process explained,\n- Provided different payment options,\n- manually processed the payment,\n- Quickpay link sent,\n- Payment completed successfully,\n- Due Date : 8/31/", "\n- confirmed the number,\n- confirmed the zip-code,\n- $4 processing fee charged,\n- updated the address,\n- provided information or transferred the call to Insurance/Assurance Department,\n- manually processed the payment,\n- guided customer to make the payment,\n- Payment completed successfully,\n- Bill Amount : 98 dollars,\n- Due Date : 8/31/", "\n- confirmed the number,\n- informed the customer to accept the terms and conditions,\n- confirmed the zip-code,\n- $4 processing fee charged,\n- Successfully restored the cancelled line,\n- informed the customer about the payment and changing billing cycle,\n- Autopay was successfully setup,\n- Quickpay link sent,\n- Payment completed successfully,\n- Due Date : 8/31/" )
期望输出
期望的输出格式如下(每列对应一类操作内容,无内容则填NA):
col1 col2 col3 col4 col5 col6 col7 col8 col9 col10 col11 col12 col13 Provided disclaimer, Disconnected NA NA NA NA NA NA NA NA NA NA NA NA NA confirmed the number pin was verified Explained reactivation confirmed the zip-code process explained changed the billing cycle Bill Amount : 60 dollars NA NA NA NA NA NA confirmed the number NA NA NA NA NA NA NA NA NA NA NA NA confirmed the number pin was verified NA NA NA NA NA NA Confirmed that the line is suspended due to non-payment NA NA NA NA confirmed the number pin was verified Explained reactivation NA NA NA NA Activation fee charged Confirmed that the line is suspended due to non-payment $4 processing fee waived informed the customer about the payment and changing billing cycle NA ...
解决方案
我们可以用tidyverse工具集里的dplyr和tidyr包来实现这个需求,步骤清晰且容易理解:
步骤1:安装并加载所需包
如果还没安装tidyverse,先执行安装命令:
install.packages("tidyverse")
加载包:
library(tidyverse)
步骤2:将数据转换为数据框
先把单列数据放到数据框中,方便后续处理:
df <- tibble(Res = extract_df$Res)
步骤3:清理并拆分每行内容
先去掉每行开头的换行符、-符号及多余空格,再将每行拆分为多个独立条目:
df_clean <- df %>% mutate( # 清理内容:移除开头的换行、-,以及条目间的多余空格 Res_clean = str_remove_all(Res, "^\\n- |\\n- |^\\n - |\\n - ") %>% # 按条目分隔符拆分,得到每个行的条目列表 str_split(",(?=\\n- )") %>% # 清理每个条目的首尾空格 map(~ str_trim(.x)) )
步骤4:转换为宽格式,相似内容归为同一列
把列表列展开成长格式,再转换为宽格式,自动将相同内容映射到同一列,缺失内容填充NA:
result <- df_clean %>% unnest_longer(Res_clean) %>% # 为每种内容类型分配唯一列名 group_by(Res_clean) %>% mutate(col_name = paste0("col", cur_group_id())) %>% ungroup() %>% # 转换为宽格式 pivot_wider( names_from = col_name, values_from = Res_clean, values_fill = NA_character_ ) %>% # 移除中间辅助列 select(-Res, -Res_clean)
步骤5:查看结果
运行View(result)或直接打印result,就能得到符合要求的多列数据,每列对应一类操作内容,无对应内容的行自动填充NA。
补充说明
- 代码会自动识别所有不同的内容类型,为每种类型创建一列,无需手动指定列名。
- 如果需要调整列的顺序,可以用
select()函数重新排列列(例如select(col2, col1, ...))。
内容的提问来源于stack exchange,提问作者george g
相关产品推荐
相关产品推荐

