鸟类观测数据按15分钟间隔拆分并按区域聚合的技术需求
Solution for Splitting Bird Observation Data into 15-Minute Bins & Aggregating by Region
Problem Naming Suggestions
Pick one based on your needs:
- Technical/Stack Overflow-focused:
R: Split Weighted Time-Series Values into Fixed 15-Minute Bins and Aggregate by Group - Bird Observation-focused:
Bird Monitoring Data: Split Corrected Duration Values into 15-Minute Intervals & Aggregate by Region
Step-by-Step R Solution
Since your sample data is an R data frame, we'll use the tidyverse ecosystem (lubridate for time handling, dplyr for data manipulation, purrr for row-wise operations) to implement your exact splitting logic.
1. Load Required Packages
library(lubridate) library(dplyr) library(purrr)
2. Define a Function to Split Single Observations into 15-Minute Bins
This function takes a single observation's start/end time and corrected value, then calculates how much of the value belongs to each overlapping 15-minute interval:
split_to_15min_bins <- function(start_time, end_time, value) { # Calculate total duration of the observation (in seconds) total_sec <- as.numeric(end_time - start_time) # Generate all 15-minute bin start times covering the observation window bin_starts <- seq( floor_date(start_time, "15 minutes"), ceiling_date(end_time, "15 minutes") - minutes(15), by = "15 minutes" ) # Iterate over each bin to calculate overlap and allocated value map_dfr(bin_starts, function(bin_start) { bin_end <- bin_start + minutes(15) # Find the overlapping time between the observation and the bin overlap_start <- max(start_time, bin_start) overlap_end <- min(end_time, bin_end) # Skip if no overlap (defensive check) overlap_sec <- as.numeric(overlap_end - overlap_start) if (overlap_sec <= 0) return(tibble()) # Allocate value based on overlap proportion allocated_val <- value * (overlap_sec / total_sec) tibble( bin_timestamp = bin_start, allocated_diff_corr = allocated_val ) }) }
3. Apply the Function to Your Data & Aggregate
# Replace `your_sample_data` with your actual data frame name final_aggregated_data <- your_sample_data %>% rowwise() %>% # Split each row into bins and store results as a list column mutate(bin_results = list(split_to_15min_bins(start, end, diff_corr))) %>% # Expand the list column into individual rows unnest(bin_results) %>% ungroup() %>% # Aggregate by region and 15-minute bin group_by(name, bin_timestamp) %>% summarize(total_corrected_duration = sum(allocated_diff_corr), .groups = "drop")
How This Matches Your Examples
- For an observation from
10:03:00-10:14:00withvalue=11: The entire window falls within the10:00:00bin, so the full 11 is allocated to that bin. - For
10:14:00-10:16:00withvalue=2: 1 minute overlaps with10:00:00, 1 minute with10:15:00→ 1 is allocated to each bin. - For
10:00:00-10:35:00withvalue=40: 15 mins in10:00, 15 mins in10:15, 5 mins in10:30→ values are40*(15/35),40*(15/35),40*(5/35)respectively.
内容的提问来源于stack exchange,提问作者Bernd V.
相关产品推荐
相关产品推荐

