You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算指定时间范围内各时段代码提交占比并定位高效工作时段

实现指定时间范围的提交时段累积占比与高效时段定位

我来帮你搞定这个时段提交占比的统计问题,结合你现有的代码和需求,我们可以一步步实现目标:

首先,先把指定时间范围的数据筛选逻辑补全(你之前的代码有点截断),确保我们只处理目标时间内的提交:

import pandas as pd

# 假设你的原始数据是self.df,先筛选时间范围
after = pd.to_datetime("2020-12-28", utc=True)
before = pd.to_datetime("2021-04-10", utc=True)
filtered_df = self.df[(self.df["date"] > after) & (self.df["date"] < before)]

接下来,我们按24小时制的小时段统计提交数(完全避开12am/pm的划分方式)。注意你的DataFrame里一个提交(同一个sha)对应多行文件记录,所以要先按提交去重,避免重复统计:

# 去重:每个提交只统计一次
unique_commits = filtered_df.drop_duplicates(subset="sha")
# 按小时统计提交量,确保按0-23的顺序排列
commit_per_hour = unique_commits["date"].dt.hour.value_counts(sort=False).sort_index()
# 重命名索引和列,让结果更清晰
commit_per_hour.index.name = "hour_of_day"
commit_per_hour.name = "commit_count"

然后计算累积占比,也就是从0点到当前小时的提交数占总提交数的百分比,这样能直观看到提交量随时间的累积趋势:

total_commits = commit_per_hour.sum()
# 计算累积提交占比,保留两位小数
cumulative_ratio = (commit_per_hour.cumsum() / total_commits * 100).round(2)
# 合并成一个DataFrame,方便查看和后续处理
hourly_stats = pd.DataFrame({
    "commit_count": commit_per_hour,
    "cumulative_percentage": cumulative_ratio
})

hourly_stats的输出格式会是这样(示例):

hour_of_daycommit_countcumulative_percentage
0123.21
185.38
.........
144562.15
.........

接下来是定位高效工作时段,我们可以通过两种方式筛选:

# 方式1:找出提交数高于平均值的时段
avg_commit = commit_per_hour.mean()
high_efficiency_hours = commit_per_hour[commit_per_hour > avg_commit].index.tolist()

# 方式2:找出提交数Top3的时段
top3_hours = commit_per_hour.nlargest(3).index.tolist()

print(f"高效工作时段(提交数高于平均):{high_efficiency_hours}")
print(f"提交数Top3时段:{top3_hours}")

如果需要可视化展示累积占比和每小时提交量,可以用matplotlib快速生成图表:

import matplotlib.pyplot as plt

plt.figure(figsize=(12,6))
# 折线图展示累积占比
plt.plot(cumulative_ratio.index, cumulative_ratio.values, marker='o', label='Cumulative Percentage')
# 柱状图展示每小时提交数
plt.bar(cumulative_ratio.index, commit_per_hour.values, alpha=0.5, label='Commit Count')
plt.xlabel('Hour of Day (24h)')
plt.ylabel('Count / Percentage')
plt.title('Hourly Commit Count & Cumulative Percentage (2020-12-28 to 2021-04-10)')
plt.legend()
plt.xticks(range(0,24))
plt.grid(axis='y', linestyle='--')
plt.show()

关键说明:

  • 我加入了drop_duplicates(subset="sha"),因为你的DataFrame里一个提交对应多行文件记录,统计提交数时应该按提交去重;如果你的需求是按文件修改数统计,可以去掉这一步。
  • 用sort_index()确保小时按0-23的自然顺序排列,避免结果混乱。
  • 累积占比用cumsum()计算后除以总数,能清晰展示提交量随时间的累积过程,方便你快速定位提交集中的时段。

内容的提问来源于stack exchange,提问作者milan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 07:03:16