You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas .duplicated()无法识别重复行的技术咨询

Troubleshooting duplicated() Not Detecting Duplicate Rows in Pandas

Let's break down why pd.DataFrame.duplicated() might not be picking up the duplicates you're seeing, and fix it step by step.

Looking at your transposed DataFrame, rows 0 and 2 (plus rows 1 and 3) appear identical at first glance—but pandas has strict rules for what counts as a duplicate, especially when dealing with NaN values and hidden data inconsistencies.

Common Causes & Fixes

1. NaN Values Are Breaking Comparisons

By default, pandas treats NaN as unequal to itself. So even if two rows have NaN in the same columns, pandas won’t consider those rows duplicates because NaN != NaN in Python. That’s likely the issue here, since your Planned_2 column is all NaNs.

Fix:

Either replace NaN with a consistent placeholder before checking duplicates, or exclude columns that are entirely NaN:

# Option 1: Fill NaNs with a unique placeholder to make comparisons work
df_filled = df.fillna("__MISSING__")
duplicate_mask = df_filled.duplicated(keep=False)
print(df[duplicate_mask])

# Option 2: Only use columns with non-NaN values for duplicate checks
valid_cols = df.columns[df.notna().any()]  # Get columns with at least one non-NaN entry
duplicate_mask = df.duplicated(subset=valid_cols, keep=False)
print(df[duplicate_mask])

2. Hidden Format/Type Inconsistencies

Even if values look identical, subtle differences can block duplicate detection:

  • String columns: Hidden whitespace (e.g., '15-07-22 ' vs '15-07-22') or mismatched date formats
  • Numeric columns: Precision differences (e.g., 68.3 vs 68.3000000001)
  • Mixed data types: A column that mixes strings and datetime objects

Fix:

Validate and clean your data first:

# Remove whitespace from string columns like Visit_Date
df['Visit_Date'] = df['Visit_Date'].str.strip()

# Round numeric columns to eliminate tiny precision gaps
df['Weight'] = df['Weight'].round(2)

# Check data types to ensure consistency across rows
print(df.dtypes)

3. Incorrect duplicated() Parameters

By default, duplicated(keep='first') only marks subsequent duplicates as True. If you want to see all duplicate rows (including the first occurrence), use keep=False:

# Show all duplicate rows, not just later ones
duplicate_mask = df.duplicated(keep=False)
print(df[duplicate_mask])

Quick Verification for Your Case

Since rows 0 and 2 match in all non-NaN columns, using the subset approach (excluding the all-NaN Planned_2 column) should correctly flag them as duplicates.

内容的提问来源于stack exchange,提问作者Hanif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:54:29