You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用函数实现Pandas Series月度数据统计与可视化(替代硬编码)

Hey there! I totally get where you're coming from—hardcoding each month is such a tedious, non-scalable hassle. Let me walk you through a couple of clean, reusable approaches to handle monthly data counting and visualization with Pandas, no more manual sum calls needed!

First, Make Sure Your Dates Are Properly Typed

Before anything else, you need to convert your string date column to a datetime type—this is critical for Pandas to recognize and group dates correctly:

import pandas as pd

# Convert your date column from string to datetime
part_date['date'] = pd.to_datetime(part_date['date'])

Method 1: Use resample() (Great for Time Series)

Pandas' resample() is built exactly for aggregating time-based data. It’s super intuitive for monthly counts:

# Resample by month ('M' = end of month; use 'MS' for start of month if preferred)
monthly_counts = part_date.resample('M', on='date').size()

# If you set 'date' as your DataFrame index first, you can simplify it to:
# part_date = part_date.set_index('date')
# monthly_counts = part_date.resample('M').size()

The result will be a Series where the index is the end-of-month date, and values are the number of records for that month.

Method 2: Use groupby() with dt.to_period()

If you prefer more explicit grouping (or want cleaner month labels like 2010-01 instead of full dates), use dt.to_period('M') to group by month-periods:

# Group by year-month periods and count records
monthly_counts = part_date.groupby(part_date['date'].dt.to_period('M')).size()

This gives you an index formatted as YYYY-MM, which is often easier to read in plots.

Visualization (Bar or Scatter Plot)

Now let’s turn those counts into a plot. Matplotlib works perfectly here, and you can easily switch between bar charts (better for comparing volumes) or scatter plots (as you mentioned):

Bar Chart (Recommended for Frequency)

import matplotlib.pyplot as plt

plt.figure(figsize=(12, 6))
monthly_counts.plot(kind='bar')
plt.title('Monthly Data Record Volume')
plt.xlabel('Month')
plt.ylabel('Number of Records')
plt.xticks(rotation=45)  # Rotate labels so they don't overlap
plt.tight_layout()  # Fix label clipping
plt.show()

Scatter Plot (As Your Original Approach)

plt.figure(figsize=(12, 6))
plt.scatter(monthly_counts.index.astype(str), monthly_counts.values)
plt.title('Monthly Data Record Volume (Scatter Plot)')
plt.xlabel('Month')
plt.ylabel('Number of Records')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()

Reusable Function for Future Use

To make this even easier, wrap everything into a reusable function—you can call it anytime with different DataFrames or date columns:

def generate_monthly_data_report(df, date_column='date', plot_type='bar'):
    # Convert date column to datetime
    df[date_column] = pd.to_datetime(df[date_column])
    # Calculate monthly counts
    monthly_counts = df.groupby(df[date_column].dt.to_period('M')).size()
    # Create plot
    plt.figure(figsize=(12, 6))
    if plot_type == 'bar':
        monthly_counts.plot(kind='bar')
    elif plot_type == 'scatter':
        plt.scatter(monthly_counts.index.astype(str), monthly_counts.values)
    # Format plot
    plt.title(f'Monthly Data Volume ({plot_type.capitalize()} Plot)')
    plt.xlabel('Month')
    plt.ylabel('Number of Records')
    plt.xticks(rotation=45)
    plt.tight_layout()
    plt.show()
    # Return the stats if you need to use them elsewhere
    return monthly_counts

# Example usage:
monthly_stats = generate_monthly_data_report(part_date, date_column='date', plot_type='scatter')

This way, you never have to write that repetitive hardcoded sum logic again—just pass in your data and go!

内容的提问来源于stack exchange,提问作者Rio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:55:43