使用pandas DataFrame列表时遇异常行为,求原因解释
Hey there! I’ve run into this exact issue before—let me break down what’s going on and how to fix it.
The Core Problem: Mutable Objects & References
Pandas DataFrames are mutable objects (just like lists in Python), which means when you pass them to a function, you’re actually passing a reference to the original object, not a copy.
If your a_function modifies the input data directly (instead of creating a new independent DataFrame), every time you call a_function(data) and append the result to list_of_df, you’re just adding another reference to the same original DataFrame to the list. By the end of the loop, all entries in list_of_df point to the exact same object—so any changes you make later (or during the loop) will show up in every "entry" in the list. That’s the "abnormal behavior" you’re seeing.
Example of the Problematic a_function
Here’s what a problematic version of your function might look like:
def a_function(df): # This modifies the original df directly df['calculated_col'] = df['some_col'] * 2 return df
Each time you call this, you’re altering the original data object, then appending its reference to the list. All entries in list_of_df are just different names for the same DataFrame.
How to Fix It: Create Copies
The fix is simple—make sure a_function returns a new copy of the DataFrame instead of modifying the original. Use the .copy() method to create an independent duplicate:
def a_function(df): # Create a full copy of the input DataFrame new_df = df.copy() # Modify the copy instead of the original new_df['calculated_col'] = new_df['some_col'] * 2 return new_df
Now every call to a_function(data) generates a brand new DataFrame, so each entry in list_of_df is a separate, independent object.
How to Verify the Issue
To confirm this is the root cause, you can print the unique identifier of each DataFrame in your loop using id():
list_of_df = [] for i in range(0,5): df = a_function(data) print(f"DataFrame {i} ID: {id(df)}") # Check if all IDs are the same list_of_df.append(df)
If all the IDs are identical, that proves every entry in the list is a reference to the same object. Once you switch to using .copy(), each ID will be unique.
内容的提问来源于stack exchange,提问作者Bravo1

