Python按月份分组datetime对象:统计月度数量并生成绘图数据
解决Python Datetime列表按年月统计(含空月份补0)的问题
哈哈,这种「看起来简单实际踩坑」的需求太有共鸣了!我之前处理类似的时间序列统计时,也差点因为忽略了空月份补0这点折腾半天。刚好帮你理清楚两种可行的解法:纯Python实现和用Pandas简化操作。
需求回顾
给定一个datetime对象列表,需要:
- X轴:覆盖从最早到最晚日期的所有年月(哪怕该月没有数据也要显示)
- Y轴:对应年月的datetime对象数量(空月份填0)
比如示例输入:[28/02/2018, 01/03/2018, 16/03/2018, 17/05/2018] → 输出(['02/2018', '03/2018', '04/2018', '05/2018'], [1, 2, 0, 1])
方法一:纯Python实现(无需第三方库)
如果不想引入Pandas,用标准库就能搞定,核心是先生成完整的年月范围,再匹配计数:
from datetime import datetime # 1. 先把输入的日期字符串转成datetime对象(如果你的输入已经是datetime可跳过这步) date_strings = ["28/02/2018", "01/03/2018", "16/03/2018", "17/05/2018"] dates = [datetime.strptime(d, "%d/%m/%Y") for d in date_strings] # 2. 统计已有月份的数量,用(year, month)元组做键(比字符串更可靠) month_counts = {} for dt in dates: key = (dt.year, dt.month) month_counts[key] = month_counts.get(key, 0) + 1 # 3. 生成从最早到最晚日期的完整年月序列 start_year, start_month = min(dates).year, min(dates).month end_year, end_month = max(dates).year, max(dates).month full_months = [] current_year, current_month = start_year, start_month while (current_year, current_month) <= (end_year, end_month): full_months.append((current_year, current_month)) # 处理跨年:12月之后跳转到次年1月 if current_month == 12: current_year += 1 current_month = 1 else: current_month += 1 # 4. 匹配计数,补全空月份的0 x_labels = [f"{month:02d}/{year}" for year, month in full_months] y_counts = [month_counts.get((year, month), 0) for year, month in full_months] print(x_labels) # 输出: ['02/2018', '03/2018', '04/2018', '05/2018'] print(y_counts) # 输出: [1, 2, 0, 1]
方法二:用Pandas简化操作(推荐)
如果你的项目已经在用Pandas,那这个需求几行代码就能搞定,resample方法会自动帮你处理空月份补0:
import pandas as pd date_strings = ["28/02/2018", "01/03/2018", "16/03/2018", "17/05/2018"] # 转成Pandas的Datetime序列 dates = pd.to_datetime(date_strings, format="%d/%m/%Y") # 按每月第一天(MS)重采样,统计数量,空月份自动补0 monthly_stats = dates.to_frame(name="count").resample("MS", on="count").count().fillna(0) # 转成需要的格式 x_labels = monthly_stats.index.strftime("%m/%Y").tolist() y_counts = monthly_stats["count"].tolist() print(x_labels) # 输出: ['02/2018', '03/2018', '04/2018', '05/2018'] print(y_counts) # 输出: [1, 2, 0, 1]
关键注意点
- 用
(year, month)元组作为统计键,比直接用字符串(比如"02/2018")更可靠,避免字符串排序时出现的逻辑错误(比如"12/2018"和"01/2019"的字符串排序结果不符合时间顺序)。 - 如果你的日期范围跨好几年,Pandas的
resample会省掉大量手动生成时间序列的代码,效率也更高。
内容的提问来源于stack exchange,提问作者EriktheRed
相关产品推荐
相关产品推荐

