Python中如何判断DataFrame是否存在重复行并返回指定提示?
Check for Duplicate Rows in DataFrame and Return Custom Message
Hey there! Here's a straightforward way to check if your DataFrame has any duplicate rows and output the exact message you need when duplicates exist:
Step-by-Step Solution
- Use
df.duplicated().any()to detect if there’s at least one duplicate row. Theduplicated()method flags each row asTrueif it’s a duplicate, andany()returnsTrueif any duplicates are present in the result. - Wrap this check in a conditional statement to print your desired message when duplicates are found.
Full Code Example
import pandas as pd # Recreate your sample DataFrame data = { 'Name': ['Jack', 'Riti', 'Aadi', 'Riti', 'Riti', 'Riti', 'Aadi', 'Sachin'], 'Age': [34, 30, 16, 30, 30, 30, 40, 30], 'City': ['Sydney', 'Delhi', 'New York', 'Delhi', 'Delhi', 'Mumbai', 'London', 'Delhi'] } df = pd.DataFrame(data) # Check for duplicates and print the target message if df.duplicated().any(): print('The df contains duplicate rows')
Quick Notes
- By default,
duplicated()marks a row as duplicate if it matches a previously seen row (usingkeep='first'). If you want to check duplicates based on specific columns instead of the entire row, use thesubsetparameter—for example:df.duplicated(subset=['Name', 'Age']).any(). - If you ever need to count the total number of duplicate rows, replace
any()withsum():df.duplicated().sum()will give you the exact count.
内容的提问来源于stack exchange,提问作者Krush23
相关产品推荐
相关产品推荐

