如何用tidyr包的separate_longer_delim()拆分带引号字符串为多行
用tidyr拆分字符串为多行的解决方案
注意事项
原始数据中002的Outcome字符串存在语法错误(Mild pain前缺失双引号),需先修正才能正常处理。
完整处理代码
library(tidyr) library(stringr) # 修正数据构造代码,补全缺失的双引号 subject <- c('001', '002') outcome <- c('["Rubbing of nose or ears","Itching of the mouth, tongue, throat","Mild pain","Eye itching"]', '["Itchy Tongue", "Mild pain","Eye irritation"]') dta <- data.frame(subject, outcome) # 清理字符串:移除首尾方括号,删除所有双引号 dta_clean <- dta %>% mutate( outcome = str_remove_all(outcome, "^\\[|\\]$"), outcome = str_remove_all(outcome, '"') ) # 按", "分隔字符串并拆分为多行 result <- dta_clean %>% separate_longer_delim(outcome, delim = ", ") # 输出结果 print(result)
处理后结果
| Subject | Outcome |
|---|---|
| 001 | Rubbing of nose or ears |
| 001 | Itching of the mouth, tongue, throat |
| 001 | Mild pain |
| 001 | Eye itching |
| 002 | Itchy Tongue |
| 002 | Mild pain |
| 002 | Eye irritation |
步骤说明
- 修正数据错误:补全
002的Outcome字符串中缺失的双引号,避免后续字符串解析异常。 - 字符串清理:用
str_remove_all移除首尾的方括号和所有双引号,将格式混乱的字符串转换为标准的,分隔文本。 - 拆分多行:使用
separate_longer_delim指定分隔符为", ",确保不会误拆分元素内部的逗号(如mouth, tongue中的逗号),将每个结果拆分为单独行。
内容的提问来源于stack exchange,提问作者biostats_charlene
相关产品推荐
相关产品推荐

