将DataFrame分组列值转换为字典的技术实现问询
Got it, let's break down how to turn that joined DataFrame into a clean, nested dictionary grouped by your job attributes. Here's a straightforward pandas-based approach that matches exactly what you're asking for.
Step 1: Set Up the Sample Data
First, let's recreate your example DataFrame so you can test the code directly:
import pandas as pd data = { 'Name': ['job1', 'job1', 'job1', 'job2', 'job2', 'job2'], 'Desc': ['desc1', 'desc1', 'desc1', 'desc2', 'desc2', 'desc2'], 'Group': ['group1', 'group1', 'group1', 'group1', 'group1', 'group1'], 'ConditionType': ['in', 'in', 'out', 'in', 'in', 'in'], 'ConditionName': ['cond1', 'cond2', 'cond1', 'cond1', 'cond2', 'cond3'] } df = pd.DataFrame(data)
Step 2: Define a Group Processing Function
We'll write a helper function to convert each job group into the dictionary structure you want. This function extracts the fixed job details and organizes conditions by their ConditionType:
def process_job_group(group): # Grab the static job details (they're the same for all rows in the group) job_details = { 'Desc': group['Desc'].iloc[0], 'Group': group['Group'].iloc[0], 'Conditions': {} } # Group conditions by type and collect names into lists condition_groups = group.groupby('ConditionType')['ConditionName'].apply(list).to_dict() job_details['Conditions'].update(condition_groups) return job_details
Step 3: Run the Group Conversion
Now group the DataFrame by Name (since each name maps to unique Desc/Group in your data) and apply the function:
final_dict = df.groupby('Name').apply(process_job_group).to_dict()
Result
The final_dict variable will be your desired nested structure:
{ 'job1': { 'Desc': 'desc1', 'Group': 'group1', 'Conditions': {'in': ['cond1', 'cond2'], 'out': ['cond1']} }, 'job2': { 'Desc': 'desc2', 'Group': 'group1', 'Conditions': {'in': ['cond1', 'cond2', 'cond3']} } }
Optional: Handle Non-Unique Name/Desc/Group Combinations
If your data ever has cases where the same Name maps to different Desc/Group values, adjust the groupby key to include all three attributes:
final_dict = df.groupby(['Name', 'Desc', 'Group']).apply( lambda g: g.groupby('ConditionType')['ConditionName'].apply(list).to_dict() ).to_dict()
This will use tuples like ('job1', 'desc1', 'group1') as keys in the top-level dictionary.
内容的提问来源于stack exchange,提问作者Max Mikhaylov

