如何基于指定范围绘制31+状态的Matplotlib Vintage Curve?
Hey there! Let's fix the issues in your current code and walk through building a proper vintage curve step by step. Your current approach isn't handling the 31+ status correctly (since those are string values like "31+") and isn't organizing data into vintage cohorts—this is the key piece missing for this type of curve.
Step 1: Clean and Prepare Your Data
First, we need to fix data types and filter for only the 31+ entries:
import numpy as np import pandas as pd import matplotlib.pyplot as plt import seaborn as sns # Load and rename columns for clarity df = pd.read_csv('sample_data.csv') df.columns = ['record_date', 'status'] # Parse dates correctly (matches your day/month/year format) df['record_date'] = pd.to_datetime(df['record_date'], format='%d/%m/%Y') # Convert status to numeric values (remove "+" signs first) df['status_numeric'] = df['status'].str.replace('+', '').astype(int) # Filter to keep only entries where status is 31+ df_31plus = df[df['status_numeric'] >= 31].copy()
Step 2: Define Vintage Cohorts
Vintage curves track groups (cohorts) of accounts by their origination month/year. For this example, let's assume record_date is when the account hit 31+ delinquency, and we need to calculate how many months after origination this happened. If you have a separate origination date column, use that instead of generating dummy data below:
# Generate dummy origination dates (replace this with your actual origination date column!) np.random.seed(42) df_31plus['origination_date'] = df_31plus['record_date'] - pd.to_timedelta(np.random.randint(1, 12, size=len(df_31plus)), unit='M') # Calculate months since origination df_31plus['months_since_origination'] = (df_31plus['record_date'].dt.year - df_31plus['origination_date'].dt.year)*12 + (df_31plus['record_date'].dt.month - df_31plus['origination_date'].dt.month) # Create vintage cohort labels (year-month of origination) df_31plus['vintage'] = df_31plus['origination_date'].dt.to_period('M').astype(str)
Step 3: Aggregate Data for the Curve
Next, we group data by vintage cohort and months since origination to count how many accounts hit 31+ status in each period:
# Aggregate counts of 31+ accounts per vintage and month since origination vintage_summary = df_31plus.groupby(['vintage', 'months_since_origination']).size().reset_index(name='count_31plus') # Optional: If you have total accounts per vintage, calculate percentage of delinquents # total_per_vintage = df.groupby('vintage').size().reset_index(name='total_accounts') # vintage_summary = vintage_summary.merge(total_per_vintage, on='vintage') # vintage_summary['percent_31plus'] = (vintage_summary['count_31plus'] / vintage_summary['total_accounts']) * 100
Step 4: Plot the Vintage Curve
Now we can visualize the curve—each line represents a different vintage cohort, showing how delinquency trends over time:
plt.figure(figsize=(12, 6)) sns.lineplot( data=vintage_summary, x='months_since_origination', y='count_31plus', # Replace with 'percent_31plus' if you calculated percentages hue='vintage', marker='o' ) plt.title('Vintage Curve: 31+ Delinquent Accounts by Origination Cohort') plt.xlabel('Months Since Account Origination') plt.ylabel('Number of 31+ Delinquent Accounts') plt.legend(title='Vintage Cohort', bbox_to_anchor=(1.05, 1), loc='upper left') plt.grid(True) plt.show()
Key Fixes from Your Original Code
- Status Handling: You tried comparing string values (like "31+") to integers—we converted status to numeric values first to fix this.
- Vintage Cohorts: You weren't grouping data by origination periods, which is essential for a vintage curve (it shows how different cohorts perform over time).
- Avoid Redundant Plots: Your loop was plotting multiple distplots unnecessarily—we aggregate data first, then plot once for clean results.
Notes for Your Specific Data
- If your
record_dateis the origination date (not delinquency date), you'll need a separate column for when the account became 31+ to calculate months since origination. - If each account has multiple status entries, first find the earliest date each account hit 31+ to avoid counting duplicates.
内容的提问来源于stack exchange,提问作者Raj

