如何对数据列的部分内容求和?15分钟流量Flashiness指数计算求助
Hey there! Let's work through how to sum specific sections of your data columns for your Flashiness Index (RB Index) project—especially since you've got multi-year data to handle. First, let's fix a tiny syntax issue in your existing code: in the summarise line, you've got an extra comma in sum(,Qi_Qi1, which will throw an error. That should be sum(Qi_Qi1, na.rm = TRUE).
Now, onto your core question: summing parts of your data (like by year, or specific time ranges). Here's how to do it with dplyr, which you're already using:
Step 1: Make sure you have a year column (if you don't already)
Assuming your dataset S has a datetime column (let's call it datetime), extract the year using lubridate (you'll need to load the package first):
library(dplyr) library(lubridate) # Add a year column to your data S <- S %>% mutate(year = year(datetime)) # Replace `datetime` with your actual time column name
Step 2: Calculate lagged flows and differences, grouped by year
To ensure you're only calculating lags within each year (so you don't get a lag from December 31 to January 1 of the next year), group by year first before creating the lagged variables:
flow_lagged_S <- S %>% group_by(year) %>% mutate( flow_lag = lag(flow, n = 1), # First-order lag within the same year Qi_Qi1 = abs(flow - flow_lag) ) %>% ungroup() # Optional: ungroup if you don't need to keep the grouping for later steps
Step 3: Sum specific sections of your data
Option 1: Calculate RB Index for each year separately
If you want the index per year, group by year again in the summarise step:
RB_index_by_year <- flow_lagged_S %>% group_by(year) %>% summarise( RB_index = sum(Qi_Qi1, na.rm = TRUE) / sum(flow, na.rm = TRUE) )
Option 2: Sum only for a specific range of years
If you need to calculate the index for a subset of years (e.g., 2015-2020), filter those rows first:
RB_index_2015_2020 <- flow_lagged_S %>% filter(year %in% 2015:2020) %>% summarise( RB_index = sum(Qi_Qi1, na.rm = TRUE) / sum(flow, na.rm = TRUE) )
Option 3: Sum for a partial year (e.g., first half of 2018)
If you want to target a specific segment within a year, add a month column and filter accordingly:
# Add month column first flow_lagged_S <- flow_lagged_S %>% mutate(month = month(datetime)) # Calculate RB Index for first half of 2018 RB_index_2018_first_half <- flow_lagged_S %>% filter(year == 2018 & month <= 6) %>% summarise( RB_index = sum(Qi_Qi1, na.rm = TRUE) / sum(flow, na.rm = TRUE) )
The key here is using group_by() to segment your data into logical groups (like years) and filter() to narrow down to the exact rows you want to sum. This keeps your calculations accurate and aligned with the time periods you care about for your Flashiness Index.
内容的提问来源于stack exchange,提问作者user9511254

