如何在DataFrame中合并两行指定列以得到目标结果?
Hey there! Let's walk through how to get your desired DataFrame result. You want to combine the two rows, keeping most values from the first row but updating specific columns (book2 and book4) with values from the second row. Here are two easy approaches:
Approach 1: Direct Column Assignment (Best for Explicit Updates)
If you know exactly which columns need to be overwritten, this is the most straightforward method. We'll directly set the values in the target row, then remove the redundant row:
import pandas as pd # Create your original DataFrame df = pd.DataFrame({ 'book1': [1, pd.NA], 'book2': [pd.NA, 'genre'], 'book3': [3, pd.NA], 'book4': [5, 1] }) # Update specific columns in row 0 with values from row 1 df.loc[0, 'book2'] = df.loc[1, 'book2'] df.loc[0, 'book4'] = df.loc[1, 'book4'] # Drop the second row and reset the index df = df.drop(index=1).reset_index(drop=True) print(df)
This will output exactly your target DataFrame. The pros here are clarity—you can see exactly which columns are being modified, no hidden logic.
Approach 2: Use combine_first() (Good for Mixed Non-Null Merges)
If you had more columns where you wanted to fill null values from the second row, plus a few to explicitly overwrite, you can use combine_first() first, then adjust the specific columns:
import pandas as pd df = pd.DataFrame({ 'book1': [1, pd.NA], 'book2': [pd.NA, 'genre'], 'book3': [3, pd.NA], 'book4': [5, 1] }) # Merge rows, filling nulls in row 0 with values from row 1 merged_row = df.loc[0].combine_first(df.loc[1]) # Overwrite the book4 value since we want row 1's value even though row 0 had a value merged_row['book4'] = df.loc[1, 'book4'] # Create the new DataFrame df = pd.DataFrame([merged_row]).reset_index(drop=True) print(df)
This method is handy if you have multiple nulls to fill across columns, and only a couple of columns that need forced updates.
内容的提问来源于stack exchange,提问作者Nima

