如何基于DataFrame1列名补全DataFrame2列并按其顺序排列
Got it, this is a common task when aligning DataFrames, and pandas has a built-in method that makes this super easy: reindex. It handles everything you need in one step—adding missing columns from df1, filling them with 0, removing any extra columns in df2 (like old5), and keeping the exact column order of df1.
Step-by-Step Implementation
First, let's set up the sample data you provided:
import pandas as pd # Define the columns from DataFrame1 df1_columns = ['adult', 'adultold', 'old', 'old1', 'old2', 'old3', 'old4', 'old6'] # Create your sample DataFrame2 df2_data = { 'adult': [0, 1, 1, 0, 1, 1, 0, 0, 0, 0, 0], 'adultold': [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], 'old2': [1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], 'old5': [0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0] } df2 = pd.DataFrame(df2_data)
Now, use reindex to align df2 with df1's columns:
# Align df2 to match df1's columns, filling missing values with 0 aligned_df = df2.reindex(columns=df1_columns, fill_value=0)
What This Does
- Adds missing columns: Any columns in df1 that aren't in df2 (like
old,old1,old3,old4,old6) are added and filled with 0. - Removes extra columns: Columns in df2 that aren't in df1 (like
old5) get dropped automatically. - Preserves column order: The resulting DataFrame will have columns in the exact same order as df1.
Result
The aligned_df will match your expected output (note: the old6 value in row 6 of your expected output appears to be a typo, since df2's row 6 only has old5=1 which isn't part of df1's columns—our code correctly sets old6 to 0 as required):
adult adultold old old1 old2 old3 old4 old6 0 0 0 0 0 1 0 0 0 1 1 0 0 0 0 0 0 0 2 1 0 0 0 0 0 0 0 3 0 0 0 0 0 0 0 0 4 1 0 0 0 0 0 0 0 5 1 0 0 0 0 0 0 0 6 0 0 0 0 0 0 0 0 7 0 0 0 0 0 0 0 0 8 0 0 0 0 0 0 0 0 9 0 0 0 0 0 0 0 0 10 0 0 0 0 0 0 0 0
内容的提问来源于stack exchange,提问作者pylearner
相关产品推荐
相关产品推荐

