如何计算指定时间范围内各时段代码提交占比并定位高效工作时段
实现指定时间范围的提交时段累积占比与高效时段定位
我来帮你搞定这个时段提交占比的统计问题,结合你现有的代码和需求,我们可以一步步实现目标:
首先,先把指定时间范围的数据筛选逻辑补全(你之前的代码有点截断),确保我们只处理目标时间内的提交:
import pandas as pd # 假设你的原始数据是self.df,先筛选时间范围 after = pd.to_datetime("2020-12-28", utc=True) before = pd.to_datetime("2021-04-10", utc=True) filtered_df = self.df[(self.df["date"] > after) & (self.df["date"] < before)]
接下来,我们按24小时制的小时段统计提交数(完全避开12am/pm的划分方式)。注意你的DataFrame里一个提交(同一个sha)对应多行文件记录,所以要先按提交去重,避免重复统计:
# 去重:每个提交只统计一次 unique_commits = filtered_df.drop_duplicates(subset="sha") # 按小时统计提交量,确保按0-23的顺序排列 commit_per_hour = unique_commits["date"].dt.hour.value_counts(sort=False).sort_index() # 重命名索引和列,让结果更清晰 commit_per_hour.index.name = "hour_of_day" commit_per_hour.name = "commit_count"
然后计算累积占比,也就是从0点到当前小时的提交数占总提交数的百分比,这样能直观看到提交量随时间的累积趋势:
total_commits = commit_per_hour.sum() # 计算累积提交占比,保留两位小数 cumulative_ratio = (commit_per_hour.cumsum() / total_commits * 100).round(2) # 合并成一个DataFrame,方便查看和后续处理 hourly_stats = pd.DataFrame({ "commit_count": commit_per_hour, "cumulative_percentage": cumulative_ratio })
hourly_stats的输出格式会是这样(示例):
| hour_of_day | commit_count | cumulative_percentage |
|---|---|---|
| 0 | 12 | 3.21 |
| 1 | 8 | 5.38 |
| ... | ... | ... |
| 14 | 45 | 62.15 |
| ... | ... | ... |
接下来是定位高效工作时段,我们可以通过两种方式筛选:
# 方式1:找出提交数高于平均值的时段 avg_commit = commit_per_hour.mean() high_efficiency_hours = commit_per_hour[commit_per_hour > avg_commit].index.tolist() # 方式2:找出提交数Top3的时段 top3_hours = commit_per_hour.nlargest(3).index.tolist() print(f"高效工作时段(提交数高于平均):{high_efficiency_hours}") print(f"提交数Top3时段:{top3_hours}")
如果需要可视化展示累积占比和每小时提交量,可以用matplotlib快速生成图表:
import matplotlib.pyplot as plt plt.figure(figsize=(12,6)) # 折线图展示累积占比 plt.plot(cumulative_ratio.index, cumulative_ratio.values, marker='o', label='Cumulative Percentage') # 柱状图展示每小时提交数 plt.bar(cumulative_ratio.index, commit_per_hour.values, alpha=0.5, label='Commit Count') plt.xlabel('Hour of Day (24h)') plt.ylabel('Count / Percentage') plt.title('Hourly Commit Count & Cumulative Percentage (2020-12-28 to 2021-04-10)') plt.legend() plt.xticks(range(0,24)) plt.grid(axis='y', linestyle='--') plt.show()
关键说明:
- 我加入了
drop_duplicates(subset="sha"),因为你的DataFrame里一个提交对应多行文件记录,统计提交数时应该按提交去重;如果你的需求是按文件修改数统计,可以去掉这一步。 - 用
sort_index()确保小时按0-23的自然顺序排列,避免结果混乱。 - 累积占比用
cumsum()计算后除以总数,能清晰展示提交量随时间的累积过程,方便你快速定位提交集中的时段。
内容的提问来源于stack exchange,提问作者milan
相关产品推荐
相关产品推荐

