如何基于现有Pandas DataFrame构造宽格式新DataFrame?
Solution to Reshape Pandas DataFrame from Long to Wide Format
First, let's recreate your sample DataFrame to work with:
import pandas as pd data = { 'codprg': [6558, 6558, 8683, 8683, 10985, 10985, 11050, 11050, 13365, 13365, 16510, 16510], 'SEX': [1.0, 2.0, 1.0, 2.0, 1.0, 2.0, 1.0, 2.0, 1.0, 2.0, 1.0, 2.0], 'counts': [9,9,9,9,9,9,4,7,9,9,8,9] } df = pd.DataFrame(data)
Now, to reshape it into your desired format (unique codprg rows, columns 1 and 2 for respective SEX counts), use Pandas' pivot() method—it's made exactly for this kind of transformation:
# Pivot the DataFrame to wide format wide_df = df.pivot(index='codprg', columns='SEX', values='counts') # Convert column names from floats (1.0, 2.0) to integers (1, 2) wide_df.columns = wide_df.columns.astype(int) # Move codprg from the index back to a regular column wide_df = wide_df.reset_index() # Ensure columns are in the exact order you requested wide_df = wide_df[['codprg', 1, 2]]
Here's the final result:
codprg 1 2 0 6558 9 9 1 8683 9 9 2 10985 9 9 3 11050 4 7 4 13365 9 9 5 16510 8 9
Quick breakdown of the steps:
pivot(...): Rearranges the data socodprgbecomes the row identifier,SEXvalues become column headers, andcountsfill the corresponding cells.columns.astype(int): Cleans up the column names to match your requested integer labels instead of floats.reset_index(): Makescodprga regular column again instead of being the DataFrame index.- Column reordering ensures the final output matches your exact column sequence.
If you ever run into duplicate codprg + SEX pairs (not present in your sample), use pivot_table() with an aggregation function like first or sum to handle duplicates:
wide_df = df.pivot_table(index='codprg', columns='SEX', values='counts', aggfunc='first').reset_index() wide_df.columns = wide_df.columns.astype(int)
内容的提问来源于stack exchange,提问作者Peter
相关产品推荐
相关产品推荐

