在R中将小时级电力负荷数据插值为半小时级数据
Here's a clean, reproducible approach using dplyr and tidyr to transform your daily hourly load data into the required half-hourly format:
Step 1: Load Required Libraries
First, make sure you have these packages installed (if not, run install.packages(c("dplyr", "tidyr"))):
library(dplyr) library(tidyr)
Step 2: Prepare the Data with Previous Day's HR24
We need the previous day's HR24 value to calculate HR0030 (the 00:30 AM reading). We'll use dplyr::lag() to fetch this:
load_demand <- load_demand %>% mutate(prev_HR24 = lag(HR24))
Step 3: Generate All Half-Hourly Columns
Now we'll create each required half-hourly column by either mapping directly to existing hourly values or calculating the linear interpolation (average, since intervals are equal) between adjacent hours:
halfhourly_load <- load_demand %>% # Calculate cross-day 00:30 reading mutate(HR0030 = (prev_HR24 + HR1)/2) %>% # Map existing hourly columns to their 00-ending half-hour names mutate(HR0100 = HR1, HR0200 = HR2, HR0300 = HR3, HR0400 = HR4, HR0500 = HR5, HR0600 = HR6, HR0700 = HR7, HR0800 = HR8, HR0900 = HR9, HR1000 = HR10, HR1100 = HR11, HR1200 = HR12, HR1300 = HR13, HR1400 = HR14, HR1500 = HR15, HR1600 = HR16, HR1700 = HR17, HR1800 = HR18, HR1900 = HR19, HR2000 = HR20, HR2100 = HR21, HR2200 = HR22, HR2300 = HR23, HR2400 = HR24) %>% # Calculate intra-day 30-minute interpolations mutate(HR0130 = (HR1 + HR2)/2, HR0230 = (HR2 + HR3)/2, HR0330 = (HR3 + HR4)/2, HR0430 = (HR4 + HR5)/2, HR0530 = (HR5 + HR6)/2, HR0630 = (HR6 + HR7)/2, HR0730 = (HR7 + HR8)/2, HR0830 = (HR8 + HR9)/2, HR0930 = (HR9 + HR10)/2, HR1030 = (HR10 + HR11)/2, HR1130 = (HR11 + HR12)/2, HR1230 = (HR12 + HR13)/2, HR1330 = (HR13 + HR14)/2, HR1430 = (HR14 + HR15)/2, HR1530 = (HR15 + HR16)/2, HR1630 = (HR16 + HR17)/2, HR1730 = (HR17 + HR18)/2, HR1830 = (HR18 + HR19)/2, HR1930 = (HR19 + HR20)/2, HR2030 = (HR20 + HR21)/2, HR2130 = (HR21 + HR22)/2, HR2230 = (HR22 + HR23)/2, HR2330 = (HR23 + HR24)/2)
Step 4: Reorder Columns to Match the Required Sequence
Finally, we'll keep only the Date column and the 48 half-hourly columns in the exact order you specified:
# Define the required column order halfhour_cols <- c("HR0030", "HR0100", "HR0130", "HR0200", "HR0230", "HR0300", "HR0330", "HR0400", "HR0430", "HR0500", "HR0530", "HR0600", "HR0630", "HR0700", "HR0730", "HR0800", "HR0830", "HR0900", "HR0930", "HR1000", "HR1030", "HR1100", "HR1130", "HR1200", "HR1230", "HR1300", "HR1330", "HR1400", "HR1430", "HR1500", "HR1530", "HR1600", "HR1630", "HR1700", "HR1730", "HR1800", "HR1830", "HR1900", "HR1930", "HR2000", "HR2030", "HR2100", "HR2130", "HR2200", "HR2230", "HR2300", "HR2330", "HR2400") # Select and reorder columns halfhourly_load <- halfhourly_load %>% select(Date, all_of(halfhour_cols))
Notes on Edge Cases
- The first row's
HR0030will beNAbecause there's no prior day'sHR24data. You can handle this by either:- Setting it to
HR1(if you assume the previous day's 24:00 is same as current day's 01:00) - Removing the first row if you don't need incomplete data
- Manually inputting the missing value if you have it
- Setting it to
Let me know if you need adjustments for specific edge cases!
内容的提问来源于stack exchange,提问作者Nick HK

