R语言按unit #2将数据框文本列拆分为两列的实现方法
R 按指定关键词拆分数据框文本列
需求说明
现有2列结构的数据框,需以关键词unit #2首次出现的位置为分界拆分Text列:
- 拆分后第2列存储
unit #2首次出现前的所有语句 - 拆分后第3列存储从
unit #2首次出现位置开始到末尾的所有语句
测试用示例数据如下:
report <- data.frame( Text = c( "unit #1 stopped at a stop sign on a road. unit #1 was speeding. unit #2 travelling southbound in lane #2 of 3 lanes. unit #2 couldn't react in time and crashed into unit #1. unit #2 was unmindful.", "unit #1 stopped there. unit #1 was under influence of drug. unit #2 travelling northbound. unit #2 was not unmindful. unit #2 crashed into unit #1.", "unit #1 was going straight. unit #1 was not speeding. unit #2 travelling southbound in lane #1 of 2 lanes. unit #2 couldn't react in time and crashed into unit #1. unit #2 was driving fast." ), id = 1:3 )
实现方法
方法1:基础R实现(无需安装额外包)
通过regexpr定位关键词首次出现的字符位置,再用字符串截取函数完成拆分,匹配时设置fixed=TRUE将关键词作为纯文本识别,避免正则特殊字符导致匹配偏差:
# 定位每条文本中unit #2首次出现的起始位置 match_pos <- regexpr("unit #2", report$Text, fixed = TRUE) # 截取前后两段文本,trimws用于去除首尾多余空格 report$text_before <- trimws(substr(report$Text, 1, match_pos - 1)) report$text_from_unit2 <- substr(report$Text, match_pos, nchar(report$Text))
方法2:stringr包实现(写法更简洁)
如果日常使用tidyverse生态的工具,可以直接用正则捕获组快速拆分:
library(stringr) library(dplyr) report <- report %>% mutate( text_before = trimws(str_extract(Text, ".*?(?=unit #2)")), text_from_unit2 = str_extract(Text, "unit #2.*") )
结果说明
拆分后以第一行数据为例:
text_before值为:unit #1 stopped at a stop sign on a road. unit #1 was speeding.text_from_unit2值为:unit #2 travelling southbound in lane #2 of 3 lanes. unit #2 couldn't react in time and crashed into unit #1. unit #2 was unmindful.
完全符合拆分要求,其余行拆分逻辑一致。
内容的提问来源于stack exchange,提问作者S Das
相关产品推荐
相关产品推荐

