基于条件批量生成命名DataFrame的Python代码优化方案求助
批量生成不同产品类型的Vintage表格优化方案
Absolutely! Your current manual approach works but gets unwieldy as you add more product types. Here's a much cleaner, scalable way to handle this using a dictionary to store all your vintage tables, paired with a simple loop:
import os import numpy as np import pandas as pd data1 = pd.read_csv('C:/Users/Oz/Desktop/vintage/vintage1.csv', encoding='latin-1') product_list = data1['product_types'].unique() def vintage_table(df): # Optional: Add df = df.copy() here if you don't want to modify the original DataFrame df['Disbursement_Date'] = pd.to_datetime(df.Disbursement_Date) df['Closing_Date'] = pd.to_datetime(df.Closing_Date) df['NPL_date'] = pd.to_datetime(df.NPL_date, errors='ignore') df['NPL_date_period'] = df.loc[df.NPL_date > '2015-01-01', 'NPL_date'].apply(lambda x: x.strftime('%Y-%m')) df['Dis_date_period'] = df.Disbursement_Date.apply(lambda x: x.strftime('%Y-%m')) df['diff'] = ((df.NPL_date - df.Disbursement_Date) / np.timedelta64(3, 'M')).round(0) df = df.groupby(['Dis_date_period','NPL_date_period']).agg({'Dis_amount' : 'sum', 'NPL_amount' : 'sum', 'diff' : 'mean'}) df.reset_index(level=0, inplace=True) df['Vintage_Ratio'] = df['NPL_amount']/df['Dis_amount'] table = pd.pivot_table(df, values='Vintage_Ratio', index='Dis_date_period', columns=['diff'],).fillna(0) return table # Batch processing with a dictionary to store results vintage_results = {} for product in product_list: # Filter data for the current product type filtered_data = data1[data1['product_types'] == product] # Generate and store the vintage table vintage_results[product] = vintage_table(filtered_data) # Access individual tables like this: # vintage_results['consumer'] # For consumer product # vintage_results['mortgage'] # For mortgage product
Why this works better:
- Scalability: No need to write new lines of code if you add more product types later—the loop handles everything automatically.
- Cleaner code: Eliminates repetitive manual assignment and reduces redundancy.
- Easy access: Retrieve any vintage table instantly using the product type name as the dictionary key.
- Maintainability: If you need to adjust how you process the data, you only modify one loop instead of multiple duplicate blocks.
A quick note: Your vintage_table function modifies the input DataFrame directly. If you want to avoid altering the original data1 dataframe, add df = df.copy() at the very start of the function—this ensures you're working with a duplicate instead of the original data.
内容的提问来源于stack exchange,提问作者ajan40
相关产品推荐
相关产品推荐

