使用pandas .duplicated()无法识别重复行的技术咨询
duplicated() Not Detecting Duplicate Rows in Pandas Let's break down why pd.DataFrame.duplicated() might not be picking up the duplicates you're seeing, and fix it step by step.
Looking at your transposed DataFrame, rows 0 and 2 (plus rows 1 and 3) appear identical at first glance—but pandas has strict rules for what counts as a duplicate, especially when dealing with NaN values and hidden data inconsistencies.
Common Causes & Fixes
1. NaN Values Are Breaking Comparisons
By default, pandas treats NaN as unequal to itself. So even if two rows have NaN in the same columns, pandas won’t consider those rows duplicates because NaN != NaN in Python. That’s likely the issue here, since your Planned_2 column is all NaNs.
Fix:
Either replace NaN with a consistent placeholder before checking duplicates, or exclude columns that are entirely NaN:
# Option 1: Fill NaNs with a unique placeholder to make comparisons work df_filled = df.fillna("__MISSING__") duplicate_mask = df_filled.duplicated(keep=False) print(df[duplicate_mask]) # Option 2: Only use columns with non-NaN values for duplicate checks valid_cols = df.columns[df.notna().any()] # Get columns with at least one non-NaN entry duplicate_mask = df.duplicated(subset=valid_cols, keep=False) print(df[duplicate_mask])
2. Hidden Format/Type Inconsistencies
Even if values look identical, subtle differences can block duplicate detection:
- String columns: Hidden whitespace (e.g.,
'15-07-22 'vs'15-07-22') or mismatched date formats - Numeric columns: Precision differences (e.g.,
68.3vs68.3000000001) - Mixed data types: A column that mixes strings and datetime objects
Fix:
Validate and clean your data first:
# Remove whitespace from string columns like Visit_Date df['Visit_Date'] = df['Visit_Date'].str.strip() # Round numeric columns to eliminate tiny precision gaps df['Weight'] = df['Weight'].round(2) # Check data types to ensure consistency across rows print(df.dtypes)
3. Incorrect duplicated() Parameters
By default, duplicated(keep='first') only marks subsequent duplicates as True. If you want to see all duplicate rows (including the first occurrence), use keep=False:
# Show all duplicate rows, not just later ones duplicate_mask = df.duplicated(keep=False) print(df[duplicate_mask])
Quick Verification for Your Case
Since rows 0 and 2 match in all non-NaN columns, using the subset approach (excluding the all-NaN Planned_2 column) should correctly flag them as duplicates.
内容的提问来源于stack exchange,提问作者Hanif

