求助:使用tidyr的gather函数将数据框转换为长格式
Hey Lisa, no need to stress—this is a super common task in pandas, and we can get it sorted quickly! Let’s walk through this with a concrete example that matches the typical wide-to-long use case, since you mentioned you’re stuck on the conversion.
Example Scenario
First, let’s assume your original wide-format DataFrame looks like this (since you didn’t share the exact content, I’ll use a relatable example):
user_id jan_sales feb_sales mar_sales 0 1 150 200 180 1 2 300 250 220 2 3 120 160 140
And you generated it with code similar to this:
import pandas as pd data = { 'user_id': [1, 2, 3], 'jan_sales': [150, 300, 120], 'feb_sales': [200, 250, 160], 'mar_sales': [180, 220, 140] } wide_df = pd.DataFrame(data)
Your target long-format DataFrame is probably structured like this, where each row represents a single user-month sales entry:
user_id month sales 0 1 jan 150 1 1 feb 200 2 1 mar 180 3 2 jan 300 4 2 feb 250 5 2 mar 220 6 3 jan 120 7 3 feb 160 8 3 mar 140
The Solution: Use pandas.melt()
The melt() function is designed exactly for this wide-to-long conversion. Here’s how to apply it to the example above:
long_df = wide_df.melt( id_vars=['user_id'], # Columns to keep as identifiers (not melted) var_name='month', # Name of the new column for the melted headers value_name='sales' # Name of the new column for the values ) # Optional: Clean up the 'month' column to remove '_sales' long_df['month'] = long_df['month'].str.replace('_sales', '')
What if Your Use Case is More Complex?
If your original DataFrame has multiple groups of columns to melt (e.g., both sales and expenses per month), you can use melt() with multiple value_vars or even pandas.wide_to_long() for more structured multi-column conversions. For example:
# If you have both sales and expenses columns wide_df_exp = pd.DataFrame({ 'user_id': [1,2,3], 'jan_sales': [150,300,120], 'jan_expenses': [50,70,40], 'feb_sales': [200,250,160], 'feb_expenses': [60,80,50] }) long_df_exp = pd.wide_to_long( wide_df_exp, stubnames=['sales', 'expenses'], i='user_id', j='month', sep='_' ).reset_index()
This will give you a clean long format with user_id, month, sales, and expenses columns.
Just adjust the id_vars, var_name, and value_name parameters to match your actual DataFrame structure, and you’ll be set!
内容的提问来源于stack exchange,提问作者Lisa Staal

