技术求助:将DataFrame转换为Array、Hashtable并按Manager迭代调整格式
First, let’s assume a sample original DataFrame structure (since you didn’t share the exact input format, this aligns with typical scenarios matching your requirements):
import pandas as pd original_df = pd.DataFrame({ 'Manager': ['Alice', 'Alice', 'Bob', 'Bob', 'Bob', 'Charlie'], 'Employee_Name': ['John', 'Jane', 'Mike', 'Sarah', 'Tom', 'Emma'], 'Position': ['Software Engineer', 'Intern', 'Data Analyst', 'Intern', 'DevOps', 'UX Designer'] })
Step 1: Add Employment Type Column
We’ll first populate the Employment_Type column using your rule: Intern roles are always Part-Time; all other positions are Full-Time. Using vectorized operations here is far more efficient than row-by-row loops:
import numpy as np original_df['Employment_Type'] = np.where(original_df['Position'] == 'Intern', 'Part-Time', 'Full-Time')
Step 2: Iterate Per Manager and Transform to Target Format
Below are two common approaches to handle per-Manager processing, depending on your exact mockup needs:
Approach 1: Process as Grouped DataFrames
If you need to work with each Manager’s data as a separate DataFrame (e.g., for exporting, reporting, or further calculations):
# Group the DataFrame by Manager manager_groups = original_df.groupby('Manager') # Iterate through each manager's group for manager_name, group_df in manager_groups: print(f"=== Data for Manager: {manager_name} ===") # Adjust columns here to match your mockup's required fields display(group_df[['Employee_Name', 'Position', 'Employment_Type']]) # Add custom logic here (e.g., save to CSV, generate summary stats)
Approach 2: Transform to Nested Structure (e.g., JSON-like dict)
If your mockup requires a nested format where each Manager maps to a list of their employees’ details:
result = {} for manager_name, group_df in manager_groups: # Convert the group to a list of dictionaries matching your mockup structure result[manager_name] = group_df[['Employee_Name', 'Position', 'Employment_Type']].to_dict('records') # Print the final structure print(result)
Step 3: Customize to Match Your Exact Mockup
If your target format has specific column names, ordering, or additional fields, tweak the code accordingly. For example, to rename columns and reorder them:
for manager_name, group_df in manager_groups: transformed = group_df.rename(columns={ 'Employee_Name': 'Full_Name', 'Position': 'Job_Role' })[['Full_Name', 'Job_Role', 'Employment_Type']] print(f"=== Transformed Data for {manager_name} ===") print(transformed)
Key Tips
- Vectorized operations (like
np.where) are critical for performance with large datasets—avoid looping through individual rows if possible. - The
groupbymethod ensures clean separation of each Manager’s data, making iteration straightforward. - If your original DataFrame has extra columns, simply adjust the subset selection (e.g.,
group_df[['Col1', 'Col2']]) to match your mockup’s required fields.
内容的提问来源于stack exchange,提问作者floader2022

