基于分类值将Pandas DataFrame多列多行合并为单行(Python)
Hey there! Let's work through this DataFrame reshaping problem you're stuck on in Python 3.4 with Pandas. Since you didn't share the exact structure of your model DataFrame or the precise desired output, I'll cover some common reshaping techniques that are likely to solve your issue, with concrete examples you can adapt.
Common Reshaping Techniques
1. Pivot Rows to Columns
If your goal is to convert categorical rows into distinct columns (e.g., grouping by an ID and spreading attributes across columns), pivot() is your go-to tool.
Sample Model Input:
import pandas as pd model = pd.DataFrame({ 'UserID': [101, 101, 102, 102], 'Metric': ['Age', 'Score', 'Age', 'Score'], 'Value': [25, 88, 30, 92] })
Code to Reshape:
# Pivot the DataFrame reshaped_df = model.pivot(index='UserID', columns='Metric', values='Value').reset_index() # Remove the auto-generated column name for cleanliness reshaped_df.columns.name = None print(reshaped_df)
Output:
UserID Age Score 0 101 25 88 1 102 30 92
2. Group and Aggregate Multiple Rows
If you need to combine multiple rows for the same group into a single row (e.g., merging lists of values), use groupby() with custom aggregation functions.
Sample Model Input:
model = pd.DataFrame({ 'OrderID': [5001, 5001, 5002, 5002, 5002], 'Product': ['Laptop', 'Mouse', 'Phone', 'Charger', 'Case'], 'Price': [999, 25, 699, 30, 15] })
Code to Reshape:
# Group by OrderID and aggregate products/prices into comma-separated strings reshaped_df = model.groupby('OrderID').agg({ 'Product': lambda x: ', '.join(x), 'Price': lambda x: ', '.join(map(str, x)) }).reset_index() print(reshaped_df)
Output:
OrderID Product Price 0 5001 Laptop, Mouse 999, 25 1 5002 Phone, Charger, Case 699, 30, 15
3. Handle Complex Reshaping with melt (If Needed)
If your desired output involves unpivoting columns into rows (the reverse of pivoting), use pd.melt(). This is useful if you need to flatten wide tables into long formats.
Next Steps
If none of these methods match your specific use case, share a small sample of your model DataFrame (you can use model.head().to_dict() to get a readable format) and the exact desired output structure. That way, we can craft a solution tailored exactly to your data.
内容的提问来源于stack exchange,提问作者Vinay Ashokkumar

