Pandas DataFrame赋值后操作同步问题解惑:为何修改副本会影响原DataFrame?
df_temp Affects df_input (And How copy() Fixes It) Great question! This is a super common "gotcha" with mutable objects in Python—including Pandas DataFrames. Let's break down exactly what's happening here:
The Problem with df_temp = df_input
When you write df_temp = df_input, you're not creating a new DataFrame. Instead, you're just assigning a new variable name that points to the exact same object in memory as df_input. Think of it like putting a second label on the same jar of cookies—if you take a cookie out using one label, the jar under the other label will also have one less cookie.
In your code, when you run df_temp.drop(index=sample.index, inplace=True), the inplace=True flag tells Pandas to modify the DataFrame directly (instead of returning a new modified copy). Since df_temp and df_input are references to the same object, this change shows up in both variables.
How df_input.copy() Solves It
Calling df_temp = df_input.copy() creates a new, independent DataFrame with its own separate space in memory. Now df_temp is a duplicate of the original df_input, so any modifications you make to df_temp (like the in-place drop) won't touch the original df_input at all.
A Quick Side Note on Shallow vs Deep Copies
By default, copy() does a "shallow copy"—this copies the DataFrame structure and references to the data values. For most cases (like the diamonds dataset, which uses basic data types like numbers and strings), this is totally fine. If your DataFrame had nested mutable objects (like lists inside cells), you'd want to use df_input.copy(deep=True) to make sure those inner objects are also copied, but that's not necessary here.
Let's Confirm with Your Code
In your original code, after the loop runs, df_input would be empty because every drop on df_temp was modifying the same underlying object. Switching to copy() ensures df_input stays intact while you manipulate df_temp to build your sampled DataFrame.
内容的提问来源于stack exchange,提问作者Bertie Bravage

