You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于条件批量生成命名DataFrame的Python代码优化方案求助

批量生成不同产品类型的Vintage表格优化方案

Absolutely! Your current manual approach works but gets unwieldy as you add more product types. Here's a much cleaner, scalable way to handle this using a dictionary to store all your vintage tables, paired with a simple loop:

import os
import numpy as np
import pandas as pd

data1 = pd.read_csv('C:/Users/Oz/Desktop/vintage/vintage1.csv', encoding='latin-1')
product_list = data1['product_types'].unique()

def vintage_table(df):
    # Optional: Add df = df.copy() here if you don't want to modify the original DataFrame
    df['Disbursement_Date'] = pd.to_datetime(df.Disbursement_Date)
    df['Closing_Date'] = pd.to_datetime(df.Closing_Date)
    df['NPL_date'] = pd.to_datetime(df.NPL_date, errors='ignore')
    df['NPL_date_period'] = df.loc[df.NPL_date > '2015-01-01', 'NPL_date'].apply(lambda x: x.strftime('%Y-%m'))
    df['Dis_date_period'] = df.Disbursement_Date.apply(lambda x: x.strftime('%Y-%m'))
    df['diff'] = ((df.NPL_date - df.Disbursement_Date) / np.timedelta64(3, 'M')).round(0)
    df = df.groupby(['Dis_date_period','NPL_date_period']).agg({'Dis_amount' : 'sum', 'NPL_amount' : 'sum', 'diff' : 'mean'})
    df.reset_index(level=0, inplace=True)
    df['Vintage_Ratio'] = df['NPL_amount']/df['Dis_amount']
    table = pd.pivot_table(df, values='Vintage_Ratio', index='Dis_date_period', columns=['diff'],).fillna(0)
    return table

# Batch processing with a dictionary to store results
vintage_results = {}
for product in product_list:
    # Filter data for the current product type
    filtered_data = data1[data1['product_types'] == product]
    # Generate and store the vintage table
    vintage_results[product] = vintage_table(filtered_data)

# Access individual tables like this:
# vintage_results['consumer']  # For consumer product
# vintage_results['mortgage']  # For mortgage product

Why this works better:

  • Scalability: No need to write new lines of code if you add more product types later—the loop handles everything automatically.
  • Cleaner code: Eliminates repetitive manual assignment and reduces redundancy.
  • Easy access: Retrieve any vintage table instantly using the product type name as the dictionary key.
  • Maintainability: If you need to adjust how you process the data, you only modify one loop instead of multiple duplicate blocks.

A quick note: Your vintage_table function modifies the input DataFrame directly. If you want to avoid altering the original data1 dataframe, add df = df.copy() at the very start of the function—this ensures you're working with a duplicate instead of the original data.

内容的提问来源于stack exchange,提问作者ajan40

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:13:19