You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas绘制按<Batch_ID,分钟>分组的唯一Code数量直方图?

问题描述

现有如下Pandas DataFrame:

import pandas as pd

data = {
    'Batch_ID': ['ABC.', 'ABC.', 'ABC.', 'ABC.', 'ABC.', 'ABB.', 'ABB.'],
    'DateTime': ['2019-01-02 17:03:41.000', '2019-01-02 17:03:41.000', 
                 '2019-01-02 17:03:42.000', '2019-01-02 17:03:48.000', 
                 '2019-01-02 17:04:41.000', '2019-01-02 17:04:41.000', 
                 '2019-01-02 17:04:45.000'],
    'Code': [230, 230, 231, 232, 230, 235, 236],
    'A1': [2., 1., 1., 2., 2., 5., 2.],
    'A2': [4, 5, 4, 7, 9, 4, 0]
}
df = pd.DataFrame(data)

需要统计每个<Batch_ID, 分钟>分组下的唯一Code数量,再生成直方图。预期统计结果如下:

<ABC, 2019-01-02 17:03> : 3
<ABC, 2019-01-02 17:04> : 1
<ABB, 2019-01-02 17:04> : 2

实现步骤

1. 处理时间格式,提取分钟级时间戳

先将DateTime列转换为Pandas可识别的datetime类型,再生成仅保留到分钟的时间列:

# 转换为datetime类型
df['DateTime'] = pd.to_datetime(df['DateTime'])
# 提取分钟级时间,格式统一为'YYYY-MM-DD HH:MM'
df['Minute_Date'] = df['DateTime'].dt.strftime('%Y-%m-%d %H:%M')

2. 分组去重并统计唯一Code数量

先对Batch_ID、Minute_Date、Code组合去重,避免同一分组内重复的Code被多次计数,再分组统计:

# 去重:保留每个分组下的唯一Code记录
unique_groups = df.drop_duplicates(subset=['Batch_ID', 'Minute_Date', 'Code'])
# 分组统计每个<Batch_ID, 分钟>下的唯一Code数量
code_count = unique_groups.groupby(['Batch_ID', 'Minute_Date'])['Code'].count().reset_index(name='Unique_Code_Count')

此时code_count的结构化结果为:

Batch_IDMinute_DateUnique_Code_Count
ABC.2019-01-02 17:033
ABC.2019-01-02 17:041
ABB.2019-01-02 17:042

3. 生成直方图

使用Matplotlib绘制直方图,将<Batch_ID, 分钟>分组作为横轴,唯一Code数量作为纵轴:

import matplotlib.pyplot as plt

# 生成符合要求的分组标签
code_count['Group_Label'] = code_count.apply(lambda row: f"<{row['Batch_ID'].rstrip('.')}, {row['Minute_Date']}>", axis=1)

# 绘制直方图
plt.figure(figsize=(10, 6))
plt.bar(code_count['Group_Label'], code_count['Unique_Code_Count'], color='skyblue')
plt.xlabel('分组(Batch_ID + 分钟级时间)')
plt.ylabel('唯一Code数量')
plt.title('各分组唯一Code数量统计')
plt.xticks(rotation=45)  # 旋转横轴标签避免重叠
plt.tight_layout()
plt.show()

4. 输出预期格式的文本结果

如果需要直接输出示例中的文本格式,执行以下代码:

for idx, row in code_count.iterrows():
    print(f"<{row['Batch_ID'].rstrip('.')}, {row['Minute_Date']}> : {row['Unique_Code_Count']}")

运行后将输出:

<ABC, 2019-01-02 17:03> : 3
<ABC, 2019-01-02 17:04> : 1
<ABB, 2019-01-02 17:04> : 2

内容的提问来源于stack exchange,提问作者Cranjis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 11:55:20