R语言提取指定行:Sequ非NA行及其对应前序NA行
提取指定行的简洁实现方案
首先给出原始数据:
df <- structure(list(Utterance = c("(5.127)", ">like I don't understand< sorry like how old's your mom¿", "(0.855)", "eh six:ty:::-one=", "(0.101)", "(0.487)", "[((v: gasps)) she said] ~no you're [not?]~", "[((v: gasps)) she said] ~no you're [not?]~", "~<[NO YOU'RE] NOT (.) you can't go !in!>~", "(0.260)", "show her [your boobs] next time"), Q = c(NA, "q_wh", "", "", NA, NA, "q_really", "", "", NA, NA), Sequ = c(NA, 1L, 1L, 1L, NA, NA, 0L, 0L, 0L, NA, NA)), class = "data.frame", row.names = c(NA, -11L))
需求说明
需要提取两类行:
- Sequ列不为NA的所有行;
- 每个连续Sequ非NA行块的前一行(该行Sequ为NA)。
最初尝试的问题
最初定义的函数仅能提取前序NA行和每个非NA块的第一行,无法获取全部目标行:
QA_sequ <- function(value) { inds <- which(!is.na(value) & lag(is.na(value))) sort(unique(c(inds-1, inds))) } library(dplyr) df %>% slice(QA_sequ(Sequ))
输出结果仅包含部分目标行,不符合预期。
简洁解决方案
直接使用dplyr的逻辑筛选,一步到位:
library(dplyr) df %>% filter( !is.na(Sequ) | (is.na(Sequ) & !is.na(lead(Sequ))) )
逻辑说明
!is.na(Sequ):直接保留所有Sequ列非NA的行;(is.na(Sequ) & !is.na(lead(Sequ))):保留那些自身Sequ为NA,但下一行Sequ非NA的行(即每个连续非NA块的前导NA行)。
执行后即可得到预期的结果:
Utterance Q Sequ 1 (5.127) <NA> NA 2 >like I don't understand< sorry like how old's your mom¿ q_wh 1 3 (0.855) 1 4 eh six:ty:::-one= 1 5 (0.487) <NA> NA 6 [((v: gasps)) she said] ~no you're [not?]~ q_really 0 7 [((v: gasps)) she said] ~no you're [not?]~ 0 8 ~<[NO YOU'RE] NOT (.) you can't go !in!>~ 0
内容的提问来源于stack exchange,提问作者Chris Ruehlemann
相关产品推荐
相关产品推荐

