Jupyter Notebook中Pandas DataFrame复制后原变量被意外修改的原因及冻结方法问询
Hey there, this isn't an environment bug—it's a classic Python/Pandas gotcha with mutable objects!
When you write temp_df = original_df, you're not making a new copy of your DataFrame. Instead, you're just giving the same underlying data a second name. So any changes you make to temp_df directly edit that shared data, which is why original_df ends up looking identical. This is how Python handles mutable objects (like DataFrames, lists, dictionaries) by default—assignments are references, not copies.
The fix is straightforward: use Pandas' .copy() method to create a separate, independent version of your DataFrame. This way, you can modify temp_df without touching the original, and you don't have to re-run the initial cell that creates original_df.
Here's what you need to do instead of temp_df = original_df:
- For a full, independent copy (this is almost always what you want to ensure original_df stays untouched):
temp_df = original_df.copy(deep=True) - If you only need a shallow copy (copies the DataFrame structure but shares nested mutable elements like lists in cells):
temp_df = original_df.copy() # Shallow is the default, so you can omit `deep=False`
Using deep=True ensures every part of the DataFrame is duplicated, so original_df stays exactly as you created it—you can keep your Jupyter Notebook split into logical, step-by-step cells for visualization and analysis as intended.
Let's walk through a simple example to confirm:
- Create your original DataFrame:
import pandas as pd original_df = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]}) - Make a proper copy:
temp_df = original_df.copy(deep=True) - Modify temp_df:
temp_df['A'] = [10, 20, 30] - Check original_df—it's still intact!
print(original_df) # Output: # A B # 0 1 4 # 1 2 5 # 2 3 6
This solves your problem completely while preserving the step-by-step workflow that makes Jupyter Notebooks so useful.
内容的提问来源于stack exchange,提问作者Christian

