基于R语言计算24小时周期内TRUE事件的发生频率
Hey there! To compute the frequency of AbovePeak = TRUE events within each 24-hour period using your dataset, here's a step-by-step approach using R's dplyr and lubridate packages:
1. Load Required Libraries
First, make sure you have these packages installed (if not, run install.packages(c("dplyr", "lubridate"))), then load them:
library(dplyr) library(lubridate)
2. Prepare Your Data
We'll start by extracting the date component from your DAT datetime column to define our 24-hour periods. By default, this uses calendar days (midnight to midnight), but I'll show how to adjust for custom windows later.
Assuming your dataset is stored in a data frame called df:
df <- df %>% mutate(calendar_date = date(DAT)) # Pulls the YYYY-MM-DD date from the datetime stamp
3. Calculate Frequency Metrics
Now group the data by year, station, and calendar date to compute key frequency stats:
- Total observations per 24-hour period
- Number of
AbovePeak = TRUEevents - Relative frequency (proportion of observations that are peak events)
daily_peak_frequency <- df %>% group_by(YEAR, STATION, calendar_date) %>% summarise( total_observations = n(), peak_event_count = sum(AbovePeak, na.rm = TRUE), peak_frequency = peak_event_count / total_observations ) %>% ungroup()
Quick Notes on the Code:
sum(AbovePeak)works because R treatsTRUEas 1 andFALSEas 0—perfect for counting events.na.rm = TRUEensures any missing values inAbovePeakdon't mess up the calculation.- Use
peak_event_countif you just want the raw number of peak events per day. Usepeak_frequencyif you want the ratio of peak events to total observations (great for irregular observation intervals).
4. Custom 24-Hour Windows (If Needed)
If your 24-hour period doesn't line up with calendar days (e.g., runs from 7 AM to 7 AM the next day), shift the datetime to group correctly:
df <- df %>% mutate( # Shift datetime back by 7 hours so 7 AM becomes midnight of our "window date" shifted_datetime = DAT - hours(7), window_date = date(shifted_datetime) ) # Now group by window_date instead of calendar_date custom_window_frequency <- df %>% group_by(YEAR, STATION, window_date) %>% summarise( total_observations = n(), peak_event_count = sum(AbovePeak, na.rm = TRUE), peak_frequency = peak_event_count / total_observations ) %>% ungroup()
Example Output
Your final data frame will look something like this (truncated):
| YEAR | STATION | calendar_date | total_observations | peak_event_count | peak_frequency |
|---|---|---|---|---|---|
| 1993 | COOP:047821 | 1993-01-01 | 8 | 1 | 0.125 |
This tells you that on January 1, 1993, at station COOP:047821, there was 1 peak event out of 8 total observations—so a 12.5% frequency of peak events that day.
内容的提问来源于stack exchange,提问作者Michaela

