Pandas绘制列累积比例及随时间变化的堆叠柱状图需求
Got it, let's tackle this step by step. You want to create a stacked bar chart showing how the relative proportions of 0, 1, and 2 change over time (using cumulative counts up to each time point), plus a reusable way to calculate those cumulative proportions. Here's exactly how to do this with Pandas and Matplotlib:
Step 1: Prepare Sample Data
First, let's build a sample DataFrame matching your example, with a time-based index (we'll use sequential daily timestamps for simplicity):
import pandas as pd import matplotlib.pyplot as plt # Your example values + time index values = [0, 1, 2, 2, 0, 0, 1] timestamps = pd.date_range(start='2024-01-01', periods=len(values), freq='D') df = pd.DataFrame({'value': values}, index=timestamps)
Step 2: Calculate Cumulative Proportions
We need to count how many times each value (0,1,2) has appeared up to each time point, then convert those counts to percentages relative to the total number of observations up to that point. This matches the example proportions you shared:
# Get cumulative counts for each value cumulative_counts = pd.DataFrame({ val: df['value'].eq(val).cumsum() for val in [0,1,2] }) # Calculate total observations up to each time point total_obs = cumulative_counts.sum(axis=1) # Convert counts to percentages (rounded to 1 decimal place) cumulative_proportions = (cumulative_counts.div(total_obs, axis=0) * 100).round(1)
If you check the output of cumulative_proportions, you'll see it matches your expected numbers:
- First row (only 0):
0: 100.0, 1: 0.0, 2: 0.0 - Fifth row (0,1,2,2,0):
0:40.0, 1:20.0, 2:40.0
Step 3: Plot the Stacked Bar Chart
Now we can use Pandas' built-in plotting to create the stacked bar chart, with customizations to make it clear and readable:
# Create stacked bar chart ax = cumulative_proportions.plot(kind='bar', stacked=True, figsize=(10,6), colormap='viridis') # Customize plot labels and styling plt.title('Cumulative Relative Proportions of 0, 1, 2 Over Time', fontsize=14) plt.xlabel('Time', fontsize=12) plt.ylabel('Relative Proportion (%)', fontsize=12) plt.xticks(rotation=45, ha='right') plt.legend(title='Value', bbox_to_anchor=(1.05, 1), loc='upper left') # Add percentage labels on each bar segment (optional but helpful) for p in ax.patches: width, height = p.get_width(), p.get_height() x, y = p.get_xy() if height > 0: # Skip labeling empty segments ax.text(x + width/2, y + height/2, f'{height}%', ha='center', va='center') plt.tight_layout() plt.show()
Step 4: Reusable Function for Cumulative Proportions
To make this easy to reuse for any similar column, wrap the calculation in a function:
def calculate_cumulative_proportions(df, column_name, target_values): """ Calculate cumulative relative proportions for specified values in a DataFrame column. Args: df (pd.DataFrame): Input DataFrame with a time-based index column_name (str): Name of the column to analyze target_values (list): List of values to calculate proportions for (e.g., [0,1,2]) Returns: pd.DataFrame: DataFrame with cumulative proportions (percentages, 1 decimal place) """ cumulative_counts = pd.DataFrame({ val: df[column_name].eq(val).cumsum() for val in target_values }) total_obs = cumulative_counts.sum(axis=1) return (cumulative_counts.div(total_obs, axis=0) * 100).round(1) # Usage example cumulative_props = calculate_cumulative_proportions(df, 'value', [0,1,2])
This function will handle the heavy lifting, so you can quickly generate proportion data for any discrete column in your DataFrame.
内容的提问来源于stack exchange,提问作者oli5679

