You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何匹配R数据框(data frame)中的对应行并完成填充清洗

R数据框清洗实现方案

依赖tidyverse套件实现,代码如下:

前置准备

library(tidyverse)

步骤1:向上回填condition列有效值

利用倒序+向下填充实现从末行有效值向上填充的需求,按participant分组处理避免跨组污染:

df_clean <- df %>%
  group_by(participant) %>%
  arrange(desc(row_number())) %>%
  fill(condition, .direction = "down") %>%
  arrange(row_number()) %>%
  ungroup()

如果Sex、Age列也是每个participant组内统一取值、仅某行有值,可以在fill函数的参数中加入这两列,一次性完成填充:

fill(Sex, Age, condition, .direction = "down")

步骤2:匹配reproduction与interval非空值

每个分组内按出现顺序给非空的interval、reproduction编号,按序号一一对应赋值:

df_clean <- df_clean %>%
  group_by(participant, condition) %>%
  mutate(
    int_idx = ifelse(!is.na(interval), cumsum(!is.na(interval)), NA),
    rep_list = list(na.omit(reproduction)),
    reproduction = ifelse(!is.na(int_idx), rep_list[[1]][int_idx], reproduction)
  ) %>%
  ungroup() %>%
  select(-int_idx, -rep_list)

步骤3:清除无效行

筛选同时有interval和reproduction有效值的行,自动保留每个分组的5行有效数据:

df_final <- df_clean %>%
  filter(!is.na(interval), !is.na(reproduction)) %>%
  arrange(participant, condition, interval)

适配说明

代码支持大规模批量处理,只要每个participant+condition分组固定有5个非空interval和5个非空reproduction值即可稳定运行,若分组逻辑有调整,修改group_by的参数即可。


内容的提问来源于stack exchange,提问作者Marc Buehner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 16:54:02