You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidyr::separate结合负向回顾断言分割字段并保留分隔符的问题

问题:使用tidyr::separate结合正则断言分割字段并保留分隔符

需求是用tidyr::separate函数分割姓名字段,需要同时处理两种情况:空格分隔的姓名(如"Another Person")和连写的大小写拼接姓名(如"HarlanNelson"),且分割后要保留姓氏的首字母,但当前实现存在两个问题:一是连写姓名分割后丢失姓氏首字母;二是无法同时适配两种分隔规则。


尝试1:仅处理连写姓名的情况

tidyr::tibble(myname = c("HarlanNelson")) |>  
  tidyr::separate(col = myname, into = c("first", "last"), sep = "(?<!^)[[:upper:]]")

输出结果:

#> # A tibble: 1 × 2
#>   first  last 
#>   <chr>  <chr>
#> 1 Harlan elson

问题:姓氏首字母"N"被当作分隔符移除,得到的姓氏是elson而非Nelson。


尝试2:同时传入空格和大写字母两种分隔规则

tidyr::tibble(myname = c("HarlanNelson", "Another Person")) |>  
  tidyr::separate(col = myname, into = c("first", "last"), sep = c(" ", "(?<!^)[[:upper:]]"))

输出结果:

#> Warning in gregexpr(pattern, x, perl = TRUE): argument 'pattern' has length > 1
#> and only the first element will be used
#> Warning: Expected 2 pieces. Missing pieces filled with `NA` in 1 rows [1].
#> # A tibble: 2 × 2
#>   first        last  
#>   <chr>        <chr> 
#> 1 HarlanNelson <NA>  
#> 2 Another      Person

问题:sep参数不支持传入多模式列表,只会使用第一个空格规则,导致连写姓名无法分割。


尝试3:加入全小写空格分隔的姓名

tidyr::tibble(myname = c("HarlanNelson", "Another Person", "someone else")) |>  
  tidyr::separate(col = myname, into = c("first", "last"), sep = c(" ", "(?<!^)[[:upper:]]"))

输出结果:

#> Warning in gregexpr(pattern, x, perl = TRUE): argument 'pattern' has length > 1
#> and only the first element will be used
#> Warning: Expected 2 pieces. Missing pieces filled with `NA` in 1 rows [1].
#> # A tibble: 3 × 2
#>   first        last  
#>   <chr>        <chr> 
#> 1 HarlanNelson <NA>  
#> 2 Another      Person
#> 3 someone      else

问题:和尝试2一致,连写姓名仍未被正确分割。


解决方案

要保留分隔符(即连写姓名的大写首字母),不能把大写字母本身作为分隔符,而是要用零宽断言匹配分割位置;同时通过|合并两种分隔规则,实现同时适配空格和连写姓名的场景。

正则逻辑说明

  • (?=\\s):匹配空格前的位置,分割空格前后的内容
  • (?<=[[:lower:]])(?=[[:upper:]]):匹配小写字母后跟大写字母的间隙,即连写姓名的分割点,不会消耗大写字母,因此能保留姓氏首字母

代码实现

tidyr::tibble(myname = c("HarlanNelson", "Another Person", "someone else")) |>  
  tidyr::separate(col = myname, into = c("first", "last"), sep = "(?=\\s)|(?<=[[:lower:]])(?=[[:upper:]])")

输出结果:

#> # A tibble: 3 × 2
#>   first  last  
#>   <chr>  <chr> 
#> 1 Harlan Nelson
#> 2 Another Person
#> 3 someone else

内容的提问来源于stack exchange,提问作者Harlan Nelson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 19:05:22