You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言使用separate_rows拆分字符串时如何保留定界符内内容不拆分

问题原因

原有双否定前瞻正则采用全局匹配逻辑,只要空格后任意位置存在定界符就不会执行拆分,当单条语句包含多个定界包裹段时,段间的普通空格会被误判为不可拆分,导致多个片段被合并。

解决方案

改用主动匹配目标片段的逻辑实现,比拆分空格的方案适配性更强,代码如下:

library(dplyr)
library(tidyr)
library(stringr)

# 构造原始数据
utt <- c("↑hey girls↑ can I <join yo:u>", "((v: grunts))", "!damn shit! got it", 
         "I mean /yeah we saw each other at a party:/↓ the other day"
)
df <- data.frame(utt)

# 拆分实现
df %>% 
  mutate(utt_split = str_extract_all(utt, "[(/≈↑£<>°!].*?[)/≈↓£<>°!]|\\S+")) %>% 
  unnest(utt_split) %>% 
  select(utt_split)

输出结果

运行上述代码得到的拆分结果与预期完全一致:

# A tibble: 14 × 1
   utt_split                         
   <chr>                             
 1 ↑hey girls↑                       
 2 can                               
 3 I                                 
 4 <join yo:u>                       
 5 ((v: grunts))                     
 6 !damn shit!                       
 7 got                               
 8 it                                
 9 I                                 
10 mean                              
11 /yeah we saw each other at a party:/↓
12 the                               
13 other                             
14 day  

逻辑说明

正则规则分为两部分,用|并联匹配:

  • 第一部分[(/≈↑£<>°!].*?[)/≈↓£<>°!]匹配所有被指定定界符包裹的内容,其中.*?为非贪婪匹配,避免相邻的两个定界段被误合并
  • 第二部分\\S+匹配所有不在定界包裹范围内的独立普通词汇

内容的提问来源于stack exchange,提问作者Chris Ruehlemann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 04:24:05