如何创建不含闰日且长度一致的Pandas DatetimeIndex
Core Idea
When generating date ranges around Feb 29 in leap years, after removing all Feb 29 timestamps, we need to pad the range with the next available timestamps from March to match the length of the same date range in non-leap years. This ensures your 3D geospatial data indexing maintains consistent window sizes across all years.
Step-by-Step Implementation
First, let's create a reusable helper function that handles the leap year adjustment automatically. It takes a base datetime (with year replaced) and your desired time window, then returns a DatetimeIndex with consistent length regardless of whether the year is a leap year:
import pandas as pd from datetime import timedelta import numpy as np def generate_consistent_date_range(base_dt, days_before=2, days_after=2, freq='6H'): # Generate the full initial time window (backward + forward) start = base_dt - timedelta(days=days_before) end = base_dt + timedelta(days=days_after) full_range = pd.date_range(start, end, freq=freq) # Filter out any timestamps falling on Feb 29 filtered_range = full_range[(full_range.day != 29) | (full_range.month != 2)] # Calculate how many timestamps we need to add to match the original length missing_count = len(full_range) - len(filtered_range) if missing_count > 0: # Generate the missing timestamps starting right after the original window ends pad_start = end + pd.Timedelta(freq) pad_range = pd.date_range(pad_start, periods=missing_count, freq=freq) # Combine, sort, and deduplicate to ensure clean ordering final_range = filtered_range.union(pad_range).sort_values() else: final_range = filtered_range return final_range
Test the Function
Let's validate with your original examples to confirm consistency:
- Leap year 2020:
test_dt = pd.to_datetime('2020-02-27 12:00:00') leap_range = generate_consistent_date_range(test_dt, days_before=0, days_after=2) print(len(leap_range)) # Output: 9 print(leap_range) # DatetimeIndex(['2020-02-27 12:00:00', '2020-02-27 18:00:00', # '2020-02-28 00:00:00', '2020-02-28 06:00:00', # '2020-02-28 12:00:00', '2020-02-28 18:00:00', # '2020-03-01 00:00:00', '2020-03-01 06:00:00', # '2020-03-01 12:00:00'], # dtype='datetime64[ns]', freq=None) - Non-leap year 2021:
test_dt = pd.to_datetime('2021-02-27 12:00:00') non_leap_range = generate_consistent_date_range(test_dt, days_before=0, days_after=2) print(len(non_leap_range)) # Output: 9
Integrate into Your Existing Code
Update your loop to use this helper function, ensuring every entry in times_list has the same length:
times = pd.date_range('1979-09-01', '1980-04-30 18:00:00', freq='6H') final_times = times[(times.day != 29) | (times.month != 2)] years = np.arange(1979, 2020, 1) times_list = [] for i in final_times: print(i) for j in years: # Replace the year and generate the consistent-length range base_dt = i.replace(year=j) total_times = generate_consistent_date_range(base_dt, days_before=2, days_after=2) times_list.append(total_times) combined_times = pd.DatetimeIndex([item for sublist in times_list for item in sublist]).sort_values()
Key Notes
- The function adapts to any frequency (not just 6H) and any time window size (adjust
days_before/days_afterto fit your analysis needs). - Padding with subsequent timestamps ensures your climate feature windows stay aligned in terms of temporal coverage, even when skipping Feb 29.
- The
unionandsort_valuessteps guarantee the final date range is ordered correctly with no duplicate entries.
内容的提问来源于stack exchange,提问作者Eli Turasky

