You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按unit #2将数据框文本列拆分为两列的实现方法

R 按指定关键词拆分数据框文本列

需求说明

现有2列结构的数据框,需以关键词unit #2首次出现的位置为分界拆分Text列:

  • 拆分后第2列存储unit #2首次出现前的所有语句
  • 拆分后第3列存储从unit #2首次出现位置开始到末尾的所有语句
    测试用示例数据如下:
report <- data.frame(
  Text = c(
    "unit #1 stopped at a stop sign on a road. unit #1 was speeding. unit #2 travelling southbound  in lane #2 of 3 lanes. unit #2 couldn't react in time  and crashed into unit #1. unit #2 was unmindful.", 
    "unit #1 stopped there. unit #1 was under influence of drug. unit #2 travelling northbound. unit #2 was not unmindful. unit #2 crashed into unit #1.", 
    "unit #1 was going straight. unit #1 was not speeding. unit #2 travelling southbound  in lane #1 of 2 lanes. unit #2 couldn't react in time and crashed into unit #1. unit #2 was driving fast."
  ), 
  id = 1:3
)

实现方法

方法1:基础R实现(无需安装额外包)

通过regexpr定位关键词首次出现的字符位置,再用字符串截取函数完成拆分,匹配时设置fixed=TRUE将关键词作为纯文本识别,避免正则特殊字符导致匹配偏差:

# 定位每条文本中unit #2首次出现的起始位置
match_pos <- regexpr("unit #2", report$Text, fixed = TRUE)

# 截取前后两段文本,trimws用于去除首尾多余空格
report$text_before <- trimws(substr(report$Text, 1, match_pos - 1))
report$text_from_unit2 <- substr(report$Text, match_pos, nchar(report$Text))

方法2:stringr包实现(写法更简洁)

如果日常使用tidyverse生态的工具,可以直接用正则捕获组快速拆分:

library(stringr)
library(dplyr)

report <- report %>%
  mutate(
    text_before = trimws(str_extract(Text, ".*?(?=unit #2)")),
    text_from_unit2 = str_extract(Text, "unit #2.*")
  )

结果说明

拆分后以第一行数据为例:

  • text_before值为:unit #1 stopped at a stop sign on a road. unit #1 was speeding.
  • text_from_unit2值为:unit #2 travelling southbound in lane #2 of 3 lanes. unit #2 couldn't react in time and crashed into unit #1. unit #2 was unmindful.
    完全符合拆分要求,其余行拆分逻辑一致。

内容的提问来源于stack exchange,提问作者S Das

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 17:54:35