You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用tidyr包的separate_longer_delim()拆分带引号字符串为多行

用tidyr拆分字符串为多行的解决方案

注意事项

原始数据中002的Outcome字符串存在语法错误(Mild pain前缺失双引号),需先修正才能正常处理。

完整处理代码

library(tidyr)
library(stringr)

# 修正数据构造代码,补全缺失的双引号
subject <- c('001', '002')
outcome <- c('["Rubbing of nose or ears","Itching of the mouth, tongue, throat","Mild pain","Eye itching"]', 
             '["Itchy Tongue", "Mild pain","Eye irritation"]')
dta <- data.frame(subject, outcome)

# 清理字符串:移除首尾方括号,删除所有双引号
dta_clean <- dta %>%
  mutate(
    outcome = str_remove_all(outcome, "^\\[|\\]$"),
    outcome = str_remove_all(outcome, '"')
  )

# 按", "分隔字符串并拆分为多行
result <- dta_clean %>%
  separate_longer_delim(outcome, delim = ", ")

# 输出结果
print(result)

处理后结果

SubjectOutcome
001Rubbing of nose or ears
001Itching of the mouth, tongue, throat
001Mild pain
001Eye itching
002Itchy Tongue
002Mild pain
002Eye irritation

步骤说明

  1. 修正数据错误:补全002的Outcome字符串中缺失的双引号,避免后续字符串解析异常。
  2. 字符串清理:用str_remove_all移除首尾的方括号和所有双引号,将格式混乱的字符串转换为标准的, 分隔文本。
  3. 拆分多行:使用separate_longer_delim指定分隔符为", ",确保不会误拆分元素内部的逗号(如mouth, tongue中的逗号),将每个结果拆分为单独行。

内容的提问来源于stack exchange,提问作者biostats_charlene

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 01:52:46