基于House列文本分割更新DataFrame并新增列的R语言实现
解决DataFrame拆分House列并新增列的问题
首先,你需要实现两个核心需求:将House列中分号分隔的内容拆分为单独行,同时新增head1、head2列。你的初始代码存在两处问题:mutate(head=head)无实际赋值逻辑,且用strsplit配合unnest处理NA值不够简洁,推荐用tidyr的separate_rows更高效。
修正后的代码
library(dplyr) library(tidyr) # 初始数据框 df <- data.frame(Region = c("AU","USA","CA","UK","GE","AU","USA","CA","UK"), lock = c(1,1,NA,1,NA,1,NA,1,NA), type= c("sale",NA,NA,"target","target",NA,"sale",NA,"target"), House =c("Tagore house","Gandhi house",NA,"Flexible;Tagore house;Gandhi house","Tagore house",NA,"Flexible;Gandhi house","Gandhi house","Gandhi house")) # 处理拆分与新增列 df_processed <- df %>% # 拆分House列分号内容为单独行,自动保留其他列对应值 separate_rows(House, sep = ";") %>% # 新增head1、head2列,示例给出两种赋值逻辑,可按需修改 mutate(head1 = case_when( grepl("Tagore", House) ~ "Tagore_group", grepl("Gandhi", House) ~ "Gandhi_group", TRUE ~ "Other" ), head2 = nchar(House)) # 用House字符长度作为示例值 print(df_processed)
代码说明
separate_rows(House, sep = ";"):直接将House列中以分号分隔的内容拆分为多行,自动同步其他列的对应值,还能妥善保留NA值的单独行。mutate部分:新增head1和head2列,示例提供了两种实用逻辑:- head1根据House包含的关键词分组;
- head2计算House的字符长度。你可以完全替换为符合业务需求的赋值逻辑(比如固定值、基于其他列的计算结果等)。
部分输出示例
Region lock type House head1 head2 1 AU 1 sale Tagore house Tagore_group 11 2 USA 1 <NA> Gandhi house Gandhi_group 12 3 CA NA <NA> <NA> Other NA 4 UK 1 target Flexible Other 8 5 UK 1 target Tagore house Tagore_group 11
内容的提问来源于stack exchange,提问作者potro
相关产品推荐
相关产品推荐

