You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr匹配数据框时间重叠行并计算对应中位数?

Solution for Calculating Median C Values Across Time Intervals

Hey there! Let's walk through how to solve this problem using dplyr—since you mentioned the between() function, we'll use that alongside other intuitive dplyr verbs to get the job done easily, even as an R beginner.

Step 1: Setup Required Packages & Data

First, make sure you have dplyr installed and loaded. If you haven't installed it yet, run this first:

install.packages("dplyr")
library(dplyr)

Next, let's replicate your original data frames so you can test the code directly:

# Your original data frames
df1 <- tibble(time1 = c(0, 1, 2, 3, 4, 5, 6, 7, 8, 9), 
              time2 = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10), 
              id = c("a", "b", "c", "d", "e", "f", "g", "h", "i", "j"))

df2 <- tibble(time = sort(runif(100, 0, 10)), 
              C = rbinom(100, 1, 0.5))

Step 2: Core Calculation with rowwise() & between()

The simplest way for beginners to handle row-by-row operations in dplyr is using rowwise(). Here's the code to add your new median column:

# Add C_median column to df1
df1_with_median <- df1 %>%
  rowwise() %>%  # Tell dplyr to process each row individually
  mutate(
    C_median = median(
      df2$C[between(df2$time, time1, time2)],  # Filter df2$C where time is in [time1, time2]
      na.rm = TRUE  # Handle cases where no rows fall in the interval (prevents NA errors)
    )
  ) %>%
  ungroup()  # Reset to normal data frame grouping after processing

How This Works

Let's break down the key parts:

  • rowwise(): This switches dplyr from operating on the entire data frame to handling each row one at a time.
  • between(df2$time, time1, time2): Creates a boolean vector where each entry is TRUE if df2$time falls between the current row's time1 and time2 (inclusive).
  • df2$C[...]: Extracts only the Cvalues fromdf2` that match the interval condition.
  • median(..., na.rm = TRUE): Calculates the median of those filtered C values. The na.rm = TRUE ensures we don't get an error if an interval has no matching rows (though with 100 rows in df2, this is unlikely here).
  • ungroup(): Always a good idea to ungroup after rowwise() to avoid unexpected behavior in future operations.

Alternative: Using purrr for Function-Style Processing

If you're curious about another approach, you can use purrr (part of the tidyverse) to iterate over rows. First load purrr:

install.packages("purrr")
library(purrr)

Then run this code:

df1_with_median <- df1 %>%
  mutate(
    C_median = map_dbl(
      1:nrow(.),  # Iterate over each row index of df1
      ~ median(df2$C[between(df2$time, time1[.x], time2[.x])], na.rm = TRUE)
    )
  )

This achieves the same result but uses functional programming instead of row-wise grouping—both are valid, so pick whichever feels more intuitive to you!

Check the Result

To verify everything worked, print the first few rows of your updated data frame:

head(df1_with_median)

You should see your original df1 columns plus a new C_median column with the median value for each interval.

内容的提问来源于stack exchange,提问作者mowglis_diaper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:29:57