You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas Timestamp绘制时段占比直方图遇重叠问题求解

问题:Pandas时间戳时段直方图条形重叠解决方法

我有一组Pandas Timestamp格式的时间戳列表:

[Timestamp('2022-01-01 21:00:00'), 
 Timestamp('2022-01-02 21:15:00'), 
 Timestamp('2022-01-03 21:00:00'), 
 Timestamp('2022-01-04 20:00:00'), 
 Timestamp('2022-01-05 21:00:00'),
 ....
 ]

想要提取时段信息,绘制以**时段(如21:00、21:15)**为标签、占比百分比为高度的直方图。尝试了以下代码,但出现条形重叠的问题:

直方图条形重叠效果

尝试的代码:

labels, counts = np.unique(histogram, return_counts=True)

all_sum = sum(counts)
percentages = [i * 100 / all_sum for i in counts]

bars = plt.bar(labels, counts, align="center", width=13, color="blue", edgecolor="black")

for i, p in enumerate(bars):
    width = p.get_width()
    height = p.get_height()
    x, y = p.get_xy()
    print(x, y)
    plt.text(x + width / 2, y + height * 1.01, "{0:.2f}".format(percentages[i]) + "%", ha="center", weight="bold")

plt.gca().set_xticks(labels)
plt.grid(False)
plt.tight_layout()

plt.show()

解决方法

重叠原因分析

你用np.unique处理后,labels是时间戳对应的原始数值(比如Pandas Timestamp转成的int64值,代表从epoch开始的纳秒数),数值跨度极大,设置的width=13完全不匹配数值尺度,导致条形直接重叠。正确思路是把时间戳转换成时分字符串作为分类标签,而非用原始时间戳数值绘图。

正确实现步骤

1. 提取时段标签

先将时间戳列表转为Pandas Series,再提取时分格式的字符串:

import pandas as pd

# 假设你的时间戳列表存在ts_list变量中
ts_series = pd.Series(ts_list)
# 提取"时:分"格式的时段标签
hour_min_labels = ts_series.dt.strftime('%H:%M')

2. 统计各时段占比

用value_counts直接统计并计算百分比:

# normalize=True直接返回占比(0-1),乘以100转为百分比
percentages = hour_min_labels.value_counts(normalize=True) * 100
# 按时段排序,让图表更整洁
percentages = percentages.sort_index()

3. 绘制直方图

此时x轴是时分字符串分类,不会有数值跨度问题,直接绘制即可:

import matplotlib.pyplot as plt

bars = plt.bar(percentages.index, percentages.values, color="blue", edgecolor="black", width=0.8)

# 为每个条形添加百分比标签
for bar in bars:
    height = bar.get_height()
    plt.text(bar.get_x() + bar.get_width()/2., height + 0.3,
             f'{height:.2f}%', ha='center', weight='bold')

# 旋转x轴标签,避免文字重叠
plt.xticks(rotation=45)
plt.ylabel('占比(%)')
plt.xlabel('时段')
plt.grid(False)
plt.tight_layout()
plt.show()

备选方案:用原始时间戳数值绘图(不推荐)

如果一定要使用np.unique得到的时间戳数值,需要将其转为分类变量:

import numpy as np

labels, counts = np.unique(hour_min_labels, return_counts=True)
percentages = counts / counts.sum() * 100

# 将数值标签转为字符串分类
plt.bar(labels.astype(str), percentages, width=0.8, color="blue", edgecolor="black")

# 后续添加标签、调整布局等操作同上

内容的提问来源于stack exchange,提问作者Luca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 22:25:18