如何用Pandas便捷查找指定日期范围内的完整周期?
在Pandas中高效筛选完全闭合周期的方案
完全可以不用大量if/else实现需求,核心思路是用Pandas的日期偏移工具生成候选周期,结合布尔筛选判断是否完全落在目标范围内,再通过字典映射替代分支判断。以下是具体实现:
步骤1:导入依赖并定义目标范围
import pandas as pd from pandas.tseries.offsets import MonthEnd, YearEnd # 示例目标范围(可替换为任意日期) start_range = pd.Timestamp("2022-12-25") end_range = pd.Timestamp("2023-02-05")
步骤2:为每种周期编写生成函数
针对不同周期类型,编写通用的生成与筛选逻辑:
月周期
def generate_months(start, end): # 生成覆盖目标范围前后各1个月的所有月初日期 month_starts = pd.date_range(start - pd.DateOffset(months=1), end + pd.DateOffset(months=1), freq='MS') periods = [] for ms in month_starts: month_end = ms + MonthEnd(1) # 筛选完全落在目标范围内的月份 if ms >= start_range and month_end <= end_range: periods.append({ 'start': ms.strftime("%Y-%m-%d"), 'end': month_end.strftime("%Y-%m-%d"), 'type': 'month', 'desc': f"month 1" }) return periods
旬周期
def generate_decades(start, end): month_starts = pd.date_range(start - pd.DateOffset(months=1), end + pd.DateOffset(months=1), freq='MS') periods = [] for ms in month_starts: # 生成当月三个旬的起始日期 decade_starts = [ms, ms + pd.DateOffset(days=10), ms + pd.DateOffset(days=20)] for idx, ds in enumerate(decade_starts, 1): # 确定旬的结束日期(第三旬取月末) decade_end = ms + MonthEnd(1) if idx == 3 else ds + pd.DateOffset(days=9) if ds >= start_range and decade_end <= end_range: periods.append({ 'start': ds.strftime("%Y-%m-%d"), 'end': decade_end.strftime("%Y-%m-%d"), 'type': 'decade', 'desc': f"decade {idx}" }) return periods
年周期
def generate_years(start, end): year_starts = pd.date_range(start - pd.DateOffset(years=1), end + pd.DateOffset(years=1), freq='YS') periods = [] for ys in year_starts: year_end = ys + YearEnd(1) if ys >= start_range and year_end <= end_range: periods.append({ 'start': ys.strftime("%Y-%m-%d"), 'end': year_end.strftime("%Y-%m-%d"), 'type': 'year', 'desc': f"year 1" }) return periods
自定义季节(匹配你的示例划分)
def generate_seasons(start, end): # 定义季节规则:起始月、结束月、名称 season_rules = [ {'start_month': 12, 'end_month': 2, 'name': 'DJF(冬季)'}, {'start_month': 3, 'end_month': 4, 'name': 'MMA(春季)'}, {'start_month': 5, 'end_month': 8, 'name': 'JJA(夏季)'}, {'start_month': 9, 'end_month': 11, 'name': 'SON(秋季)'} ] periods = [] # 覆盖目标范围前后1年,处理跨年度季节 for year in range(start.year - 1, end.year + 1): for idx, rule in enumerate(season_rules, 1): s_month, e_month = rule['start_month'], rule['end_month'] # 处理跨年度季节(如DJF为前一年12月至当年2月) if s_month > e_month: season_start = pd.Timestamp(f"{year-1}-{s_month}-01") season_end = pd.Timestamp(f"{year}-{e_month}-01") + MonthEnd(1) else: season_start = pd.Timestamp(f"{year}-{s_month}-01") season_end = pd.Timestamp(f"{year}-{e_month}-01") + MonthEnd(1) if season_start >= start_range and season_end <= end_range: periods.append({ 'start': season_start.strftime("%Y-%m-%d"), 'end': season_end.strftime("%Y-%m-%d"), 'type': 'season', 'desc': f"season {idx}({rule['name']})" }) return periods
步骤3:用字典映射替代if/else
将周期类型与对应生成函数绑定,无需分支判断即可扩展新周期:
# 映射周期类型到生成函数 period_map = { 'month': generate_months, 'decade': generate_decades, 'year': generate_years, 'season': generate_seasons } # 用户指定要查询的周期类型 target_periods = ['month', 'decade'] # 收集所有符合条件的周期 final_results = [] for p_type in target_periods: final_results.extend(period_map[p_type](start_range, end_range)) # 打印结果(或转为DataFrame查看) for res in final_results: print(f"{res['start']} 至 {res['end']}, {res['type']} {res['desc']}")
效果验证
针对你给出的第一个示例,运行后会输出:
2023-01-01 至 2023-01-31, month month 1 2023-01-01 至 2023-01-10, decade decade 1 2023-01-11 至 2023-01-20, decade decade 2 2023-01-21 至 2023-01-31, decade decade 3
若切换到第二个示例的范围(start_range = pd.Timestamp("2021-12-31"), end_range = pd.Timestamp("2023-02-28")),并指定target_periods = ['year', 'season'],输出将完全匹配你给出的结果。
这种方案的优势:
- 新增周期类型仅需添加生成函数和字典条目,无需修改核心逻辑
- 利用Pandas内置日期偏移自动处理月末/年末的日期计算,避免手动判断天数
- 通过扩展生成范围+精准筛选,确保不会遗漏或错误匹配周期
内容的提问来源于stack exchange,提问作者Droid
相关产品推荐
相关产品推荐

