如何在Python函数中过滤DataFrame数据?过滤环节报错求助
Hey there, let's break down what's causing issues with your function and get it working smoothly. The main problem here is likely related to pandas view vs copy behavior when you filter the DataFrame, plus some potential pitfalls with in-place modifications.
First, let's look at the root of the error: when you run df = df[df['NEW'].isin(LF)], you're creating a slice of the original DataFrame. In some cases, this is just a view (not an independent copy), so when you try to modify columns later like df[AN] = "0000" + ..., you might hit a SettingWithCopyWarning or even a hard error because pandas isn't sure if you want to modify the original data or the slice.
Here's the revised function with fixes and explanations:
import pandas as pd def Acc(df, AN, TR, LF): # Create a full copy of the input DataFrame to avoid modifying the original # and prevent view/copy confusion working_df = df.copy() # Convert columns to strings safely working_df[AN] = working_df[AN].astype(str) working_df[TR] = working_df[TR].astype(str) # Calculate the string length for filtering working_df['NEW'] = working_df[AN].str.len() # Filter the data - this creates a new DataFrame (not a view) filtered_df = working_df[working_df['NEW'].isin(LF)] # Use .loc to safely assign the new value to the AN column # This ensures we're modifying the actual DataFrame, not a view filtered_df.loc[:, AN] = "0000" + filtered_df[TR] + "/" + filtered_df[AN] # Optional: Drop the temporary 'NEW' column if you don't need it in the output filtered_df = filtered_df.drop(columns=['NEW']) return filtered_df
Key Improvements:
- Copy the input DataFrame: Using
.copy()ensures we're working on an independent dataset, so we don't accidentally alter the originalraw_fileand avoid view/copy issues. - Explicit filtered DataFrame: Separating the filtered data into its own variable makes the logic clearer and avoids overwriting objects in-place.
.locfor column assignment: This is the safest way to modify columns in pandas, as it explicitly targets the DataFrame's data instead of a potential view, eliminating warnings/errors from ambiguous assignments.
Testing with Your Sample Data:
If you run this with your sample raw_file:
raw_file = pd.DataFrame({ 'NAME': ['abc', 'def', 'ghi', 'jkl', 'mno'], 'ABCD': [1, 254, 8976541, '000000111/1215', 15614987], 'XYZ': [111, 121, 254, 111, 117] }) # Call the function result = Acc(raw_file, AN='ABCD', TR='XYZ', LF=[1,2,3,4,14])
The filtered result will keep rows where ABCD string length is in [1,2,3,4,14]:
abc(length 1) → becomes0000111/1def(length 3) → becomes0000121/254
If you still run into issues, check these quick fixes:
- Ensure
LFis a list of integers (sinceNEWholds integer lengths) - Handle missing values: If your
ANcolumn has NaNs, addworking_df[AN] = working_df[AN].fillna('').astype(str)to avoid'nan'strings (length 3) being incorrectly filtered.
内容的提问来源于stack exchange,提问作者Aditya

