Pandas DataFrame.update()中overwrite参数作用及差异示例咨询
overwrite Parameter Let me break down exactly what the overwrite parameter does, and show you a clear example where setting it to True vs False makes a noticeable difference.
First, the core purpose of overwrite:
- When
overwrite=True(this is the default), any non-NA value in the right-hand DataFrame will replace the corresponding value in the original DataFrame—even if the original value is not NA. - When
overwrite=False, the right-hand DataFrame's non-NA values will only fill in NA values in the original DataFrame; existing non-NA values in the original will stay untouched.
The reason you might not have seen a difference in your tests is probably because your test cases didn't hit the key scenario: where the original DataFrame has non-NA values in positions that the right-hand DataFrame also has non-NA values for. Let's fix that with a concrete example.
Example with Clear Differences
Let's create our original DataFrame first:
import pandas as pd original_df = pd.DataFrame({ 'A': [1, 2, None], 'B': [10, None, 30], 'C': [100, 200, 300] })
This gives us:
A B C 0 1.0 10.0 100 1 2.0 NaN 200 2 NaN 30.0 300
Now create the update DataFrame—note that it has non-NA values in positions where original_df already has non-NA values (like column A, row 0; column C, row 1):
update_df = pd.DataFrame({ 'A': [99, None, 98], 'B': [None, 25, None], 'C': [None, 299, 399] })
Which looks like:
A B C 0 99.0 NaN NaN 1 NaN 25.0 299.0 2 98.0 NaN 399.0
Case 1: overwrite=True (Default)
When we run original_df.update(update_df, overwrite=True):
original_df_copy1 = original_df.copy() original_df_copy1.update(update_df, overwrite=True) print(original_df_copy1)
Output:
A B C 0 99.0 10.0 100.0 1 2.0 25.0 299.0 2 98.0 30.0 399.0
Notice here:
- Column A, row 0: original
1.0is replaced with99.0 - Column C, row 1: original
200is replaced with299.0 - Column C, row 2: original
300is replaced with399.0 - NA values in original (like B row1, A row2) are filled as expected.
Case 2: overwrite=False
Now run with overwrite=False:
original_df_copy2 = original_df.copy() original_df_copy2.update(update_df, overwrite=False) print(original_df_copy2)
Output:
A B C 0 1.0 10.0 100 1 2.0 25.0 200 2 98.0 30.0 300
Here's the key difference:
- Column A, row 0: original
1.0stays (not overwritten by99.0) - Column C, row1: original
200stays (not overwritten by299.0) - Column C, row2: original
300stays (not overwritten by399.0) - Only the NA values in original (B row1, A row2) are filled with values from
update_df.
That's the critical distinction! The overwrite parameter controls whether existing non-NA values in your original DataFrame get replaced by non-NA values from the update DataFrame.
内容的提问来源于stack exchange,提问作者Zhang18

