You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于非规整R数据集拆分District至新列并匹配分组?

处理非规整数据集:提取District至单独列并匹配对应行

我正在处理一个非规整数据集,其中location列混有包含"District"的行(如"District 1"、"District 2"等),需要将这些District内容提取至单独列,并与下方对应的location行匹配。

数据集片段

location     iri length
   <chr>      <dbl>  <dbl>
 1 District 1 NA        NA
 2 S00180CB    4.60  32171
 3 S00186CB    5.17  31286
 4 S00188CB    6.02   4742
 5 District 2 NA        NA
 6 S00136CB    4.91   5968
 7 S00285CB    4.33  29340
 8 S00288CB   10.9     141
 9 S00289CB    4.41   9126
10 District 3 NA        NA
11 S00231CB    4.34   6895
12 S00266CB    5.65  18985
13 S00381CB    4.39  13799
14 S00382CB    8.96    124
15 District 4 NA        NA
16 S00303CB    4.11  79599
17 S00311CB    3.17    625
18 District 5 NA        NA
19 S00163CB    3.49  17996
20 S00253CB    2.65    905
21 S00259CB    2.26    905

尝试的代码

library(tidyverse)

df |> 
  mutate(district = ifelse(str_detect(location, "District"), 
                           rep_len("District 1", length.out = n()), "none")

预期输出

district   location   iri length
   <chr>      <chr>    <dbl>  <dbl>
 1 District 1 S00180CB  4.60  32171
 2 District 1 S00186CB  5.17  31286
 3 District 1 S00188CB  6.02   4742
 4 District 2 S00136CB  4.91   5968
 5 District 2 S00285CB  4.33  29340
 6 District 2 S00288CB 10.9     141
 7 District 2 S00289CB  4.41   9126
 8 District 3 S00231CB  4.34   6895
 9 District 3 S00266CB  5.65  18985
10 District 3 S00381CB  4.39  13799
11 District 3 S00382CB  8.96    124
12 District 4 S00303CB  4.11  79599
13 District 4 S00311CB  3.17    625
14 District 5 S00163CB  3.49  17996
15 District 5 S00253CB  2.65    905
16 District 5 S00259CB  2.26    905

但实际结果不符合预期,求更高效的实现方法。


解决方案

可以利用tidyverse中的fill()函数实现需求,步骤清晰且高效:

library(tidyverse)

df_clean <- df |>
  # 生成district列:匹配到District行时赋值,否则为NA
  mutate(district = if_else(str_detect(location, "District"), location, NA_character_)) |>
  # 向下填充NA值,让后续行继承上方的District分组
  fill(district, .direction = "down") |>
  # 过滤掉原始的District标记行
  filter(!str_detect(location, "District")) |>
  # 调整列顺序,与预期输出一致
  select(district, everything())

代码说明

  1. mutate()创建分组列:通过str_detect()识别District行,将对应值存入district列,其他行设为NA。
  2. fill()填充分组值:使用向下填充逻辑,让每个District行下方的所有数据行自动匹配对应的分组。
  3. filter()清理冗余行:移除作为分组标记的District行,只保留有效数据。
  4. select()调整列顺序:将district列移至首位,对齐预期输出格式。

这个方法无需手动指定分组值,能自动适配任意数量的District分组,完全满足需求。


内容的提问来源于stack exchange,提问作者dfc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 03:50:19