不修改原数据集的前提下,用其他列值填充Pandas DataFrame空单元格的更优实现方式
copy() Needed)? I have a Pandas DataFrame structured like this:
| Company Name | Person Name | Phone number |
|---|---|---|
| General Electric | John Doe | |
| Ford | 123 456 789 |
I need to fill the null values in the Person Name column with the corresponding values from the Company Name column. The desired result is:
| Company Name | Person Name | Phone number |
|---|---|---|
| General Electric | John Doe | |
| Ford | Ford | 123 456 789 |
I can use this code to achieve the fill:
df.loc[df["Person Name"].isna(),'Person Name'] = df["Company Name"]
But this modifies the original DataFrame. Using df.copy().loc[df["Person Name"].isna(),'Person Name'] = df["Company Name"] works without altering the original data, but I'm wondering if there's a more elegant way to do this without using copy()?
Great question! You have two clean, idiomatic Pandas approaches that avoid copy() entirely, both returning a new DataFrame without touching your original data:
1. Use Series.fillna() with DataFrame.assign()
This method explicitly creates a new version of the Person Name column in a fresh DataFrame, leaving your source data untouched:
new_df = df.assign(Person_Name=df["Person Name"].fillna(df["Company Name"]))
assign() always returns a new DataFrame, and fillna() here replaces nulls in Person Name with matching Company Name values—no modifications to the original df whatsoever.
2. Use DataFrame.fillna() with a dictionary parameter
If you prefer a concise one-liner (especially useful if you need to fill multiple columns at once), pass a dictionary to fillna() that maps column names to their fill values:
new_df = df.fillna({"Person Name": df["Company Name"]})
This also generates a brand new DataFrame, keeping your original dataset completely intact.
Why these are better than copy().loc
- Readability: These methods clearly signal your intent (filling nulls without modifying source data) to anyone reading your code.
- Idiomatic Pandas:
fillna()andassign()are designed specifically for this kind of immutable data manipulation, aligning with Pandas' best practices. - Efficiency: No explicit
copy()call means you avoid unnecessary memory overhead (though minimal, it’s still cleaner and more intentional).
内容的提问来源于stack exchange,提问作者Noan Cloarec

