如何在保留date变量的前提下重塑含2列的DataFrame?
Hey there! I totally get how frustrating it is when you’re trying to reshape your DataFrame and end up losing the date column—especially when most examples out there use a separate date column structure that doesn’t match yours. Let’s walk through simple, reliable ways to keep that date intact no matter what reshaping you’re doing.
First, Let’s Define a Common Scenario
Let’s start with a sample DataFrame that mirrors your setup (one single date column plus other values you want to reshape):
import pandas as pd # Your original DataFrame structure df = pd.DataFrame({ 'date': ['2023-01-01', '2023-01-02', '2023-01-03'], 'product_a': [10, 15, 20], 'product_b': [5, 8, 12], 'product_c': [3, 4, 6] })
1. Reshaping Wide → Long (Keeping Date)
If you’re converting a wide table to a long one, the key is to specify date as an identifier variable with melt(). This tells pandas to keep date as a column while unpivoting the rest:
# Convert to long format, preserve date long_df = df.melt( id_vars='date', # This keeps date from being unpivoted var_name='product', # Name for the new column of original headers value_name='sales' # Name for the new column of values )
Result snippet:
| date | product | sales |
|---|---|---|
| 2023-01-01 | product_a | 10 |
| 2023-01-02 | product_a | 15 |
2. Reshaping Long → Wide (Keeping Date)
If you’re going from long to wide, use pivot() and set date as the index (then reset it to move it back to a column if needed):
# Sample long-format DataFrame long_df = pd.DataFrame({ 'date': ['2023-01-01', '2023-01-01', '2023-01-02', '2023-01-02'], 'product': ['a', 'b', 'a', 'b'], 'sales': [10, 5, 15, 8] }) # Convert to wide format, keep date as a column wide_df = long_df.pivot( index='date', # Use date as the row identifier columns='product', # Columns come from product names values='sales' # Values fill the table ).reset_index() # Move date back from index to a regular column # Clean up column names (optional) wide_df.columns.name = None
Result snippet:
| date | a | b |
|---|---|---|
| 2023-01-01 | 10 | 5 |
| 2023-01-02 | 15 | 8 |
3. For More Complex Reshaping (Stack/Unstack)
If you’re dealing with multi-level columns or indexes, make sure date is part of the index before using stack() or unstack():
# Set date as the index first df_indexed = df.set_index('date') # Stack columns into rows (preserves date as index) stacked_df = df_indexed.stack().reset_index() stacked_df.columns = ['date', 'product', 'sales'] # Rename columns for clarity
The Core Rule to Remember
No matter which reshaping method you use: always explicitly include date as an identifier (either via id_vars for melt, index for pivot, or part of the index for stack/unstack). This prevents pandas from treating it as a value to reshape and losing it in the process.
内容的提问来源于stack exchange,提问作者Niccola Tartaglia

