如何按组大小将信用卡交易分组数据划分为四分位数并导出CSV
Got it, let's walk through how to split your grouped credit card transactions into quartiles by group size and save each group to the correct directory structure. Here's a step-by-step solution that builds on your existing code:
First, we need to assign each group to a quartile based on its size. We'll use pd.qcut to split groups into 4 equal-sized buckets (or use pd.cut if you prefer fixed value ranges instead of equal counts).
import pandas as pd import os # Your existing code to load and group data # raw_data = pd.read_csv(dataset_path) raw_data = pd.read_csv('https://query.data.world/s/xapduodvxrzul6tjrupmlqiot4bpex') categories = ('Cardholder Last Name', 'Cardholder First Initial') groups = raw_data.groupby(list(categories)) size_of_groups = groups.size() # Assign quartile labels to each group quartile_labels = ['1st_quartile', '2nd_quartile', '3rd_quartile', '4th_quartile'] # qcut splits groups into 4 equal-count buckets group_quartiles = pd.qcut(size_of_groups, q=4, labels=quartile_labels) # Combine group sizes and quartile labels for easy lookup group_metadata = pd.DataFrame({'group_size': size_of_groups, 'quartile': group_quartiles})
We'll use Python's os module to automatically create the required folder hierarchy—no manual folder setup needed:
root_directory = "Groups" # Create each quartile subdirectory (ignore if already exists) for quartile in quartile_labels: target_dir = os.path.join(root_directory, quartile) os.makedirs(target_dir, exist_ok=True)
Loop through each grouped dataset, look up its assigned quartile, and save it to the corresponding folder with the requested filename format:
for group_key, group_data in groups: # Get the quartile for the current group assigned_quartile = group_metadata.loc[group_key, 'quartile'] # Generate filename from the group key (e.g., ('Mason', 'C') becomes "(Mason, C).csv") filename = f"{group_key}.csv" # Build full save path save_path = os.path.join(root_directory, assigned_quartile, filename) # Save the group data to CSV (exclude index for cleaner output) group_data.to_csv(save_path, index=False)
If you want to split groups by fixed size ranges instead of equal counts (e.g., groups with size ≤ Q1 go to 1st quartile), use pd.cut instead of pd.qcut:
# Calculate fixed quartile thresholds q1 = size_of_groups.quantile(0.25) q2 = size_of_groups.quantile(0.5) q3 = size_of_groups.quantile(0.75) # Assign quartiles based on value ranges group_quartiles = pd.cut( size_of_groups, bins=[-float('inf'), q1, q2, q3, float('inf')], labels=quartile_labels, include_lowest=True )
To check if groups are distributed as expected, run this to see the count of groups per quartile:
print(group_metadata['quartile'].value_counts())
内容的提问来源于stack exchange,提问作者user13051721

