You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于House列文本分割更新DataFrame并新增列的R语言实现

解决DataFrame拆分House列并新增列的问题

首先,你需要实现两个核心需求:将House列中分号分隔的内容拆分为单独行,同时新增head1、head2列。你的初始代码存在两处问题:mutate(head=head)无实际赋值逻辑,且用strsplit配合unnest处理NA值不够简洁,推荐用tidyr的separate_rows更高效。

修正后的代码

library(dplyr)
library(tidyr)

# 初始数据框
df <-  data.frame(Region = c("AU","USA","CA","UK","GE","AU","USA","CA","UK"),
                  lock = c(1,1,NA,1,NA,1,NA,1,NA),
                  type= c("sale",NA,NA,"target","target",NA,"sale",NA,"target"),
                  House =c("Tagore house","Gandhi house",NA,"Flexible;Tagore house;Gandhi house","Tagore house",NA,"Flexible;Gandhi house","Gandhi house","Gandhi house"))

# 处理拆分与新增列
df_processed <- df %>%
  # 拆分House列分号内容为单独行,自动保留其他列对应值
  separate_rows(House, sep = ";") %>%
  # 新增head1、head2列,示例给出两种赋值逻辑,可按需修改
  mutate(head1 = case_when(
           grepl("Tagore", House) ~ "Tagore_group",
           grepl("Gandhi", House) ~ "Gandhi_group",
           TRUE ~ "Other"
         ),
         head2 = nchar(House)) # 用House字符长度作为示例值

print(df_processed)

代码说明

  • separate_rows(House, sep = ";"):直接将House列中以分号分隔的内容拆分为多行,自动同步其他列的对应值,还能妥善保留NA值的单独行。
  • mutate部分:新增head1和head2列,示例提供了两种实用逻辑:
    • head1根据House包含的关键词分组;
    • head2计算House的字符长度。你可以完全替换为符合业务需求的赋值逻辑(比如固定值、基于其他列的计算结果等)。

部分输出示例

Region lock   type         House       head1 head2
1      AU    1   sale  Tagore house Tagore_group    11
2     USA    1   <NA>  Gandhi house Gandhi_group    12
3      CA   NA   <NA>          <NA>        Other    NA
4      UK    1 target      Flexible        Other     8
5      UK    1 target  Tagore house Tagore_group    11

内容的提问来源于stack exchange,提问作者potro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 11:43:25